Repository navigation
feat(blog): split brand from writing context, and enrich both - #7623
Open
aka-sacci-ccr wants to merge 47 commits into
Open
aka-sacci-ccr wants to merge 47 commits into
aka-sacci-ccr wants to merge 47 commits into
Conversation
The blog's editorial context was one JSON block and one AI pass. Two things were tangled in it: who the brand is, which holds whether or not it ever publishes a post, and how the blog is written, which belongs to the blog and evolves with it. Split them into `blog-manager-brand` and `blog-manager-context`, inferred by two chained tools — BLOG_CONTEXT_EXTRACT takes BLOG_BRAND_EXTRACT's answer, because writing rules read out of a site's copy are only as good as the understanding of whose copy it is. Sites written before the split keep working: the writing rules fall back to the brand block until the new one exists, and each save writes only the fields its block owns, so the old block sheds them. Six new fields — special dates, keywords, differentiators and commercial policies on the brand; vocabulary and example phrases on the generation side. Example phrases carry a sounds/doesn't-sound toggle, since a counter-example is what pins down how far the imitable ones go. Three new evidence sources feed the same one button: - SEO, which was never read on purpose. The site-level default lives on `site/apps/site.ts`, so no evidence tier reached it, and page SEO leaked in unlabelled. It is a site's most deliberately written copy. - The store's catalog, through the existing catalog-invoke proxy. The VTEX account stays in the running site's own app; nothing new is fetched cross-origin. - Web research, broadened past competitors to dates and differentiators — the three things a site structurally cannot say about itself. Its citations are captured so a human can check an invented date. Pages now rank home, then institutional, then commerce: a PDP is a template with a product name substituted in, so the thousandth teaches nothing the first did not. Also drops the "coming soon" veil over the Context tab and the disabled state on idea and post generation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # apps/web/src/components/sandbox/content/blog/use-generate-post.ts # apps/web/src/i18n/en/common.ts # apps/web/src/i18n/pt-br/common.ts
# Conflicts: # apps/web/src/components/sandbox/content/blog/blog-data.test.ts
# Conflicts: # apps/web/src/components/sandbox/content/blog/posts-workspace.tsx
…elds
The brand extract filled fields it had no evidence for. Asked for commercial
policies on a site that publishes none, it returned that the site "references
'Frete Grátis' in its metadata" — true about the metadata, useless as a policy,
and indistinguishable from a real one once persisted. Every post generated
afterwards inherits it.
The extractor cannot be the judge of its own output, so a second pass reads the
same evidence and scores each extracted claim on two axes that come apart:
confidence (is it true?) and relevance (would a writer do anything differently
because of it?). That line scores high on the first and near zero on the second,
which is why one number could not have caught it. A claim must reach 85 on both
to be written; the rest are counted in the toast and detailed in the server log.
Shaped after tools/task-board/duplicate-check.ts: a model verdict, a pure gate
over it, and every failure path reading as the lenient outcome. Coverage is all
or nothing — a judge that answers for 30 of 45 claims has truncated, not
rejected, so partial coverage keeps everything and reports judged: false. The
gate must never be the reason Preencher fills nothing.
Also:
- Remove `differentiators`. Neither the site nor the web answers "what a rival
cannot claim", so the model invented it; `competitors` covers the positioning
with real evidence. Sites that saved the field shed it on the next autosave,
the way blog-data.ts already documents for legacy fields — anyone who wrote
differentiators by hand loses them.
- `keywords` becomes a plain string[], edited as chips. It needs its own
normalize path: normalizeBrandRules tolerates bare strings and maps them back
to { name, value }, so leaving keywords in the rule-field list would have
re-objectified them on the next autosave with no error anywhere.
- Only the form obeys the gate. BLOG_CONTEXT_EXTRACT is told to write in the
brand's `language`, so a cut scalar falls back to what is already written
rather than setting the whole writing pass adrift.
- voiceExamples are checked for verbatim presence before the judge sees them:
the schema demands exact quotes, which makes that a substring test rather than
an opinion.
Cost: 4 model calls per fill becomes 6.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gate was cutting good extractions. Fourteen discards on one run, and the pattern in the log is one mistake made three ways. Relevance was asked globally — "would a writer do anything differently because of this?" — and almost nothing clears that bar on its own. Five voice examples with confidence 98, verbatim and verified, were cut at relevance 65-75. But a voice example is not supposed to change what a post says; it is supposed to show how the brand sounds, and those did. Same for vocabulary: "Ar Indireto" at 90/80 is exactly what that field exists to capture. So relevance is now asked per field. Every FieldSpec carries the question relevance means for it, and the judge gets them as a legend keyed by field name alongside the claims. The question for competitors is "is this a real rival, with something said about how it differs?", not "is this useful?". Also: - `companyName` and `language` leave the gate entirely. They are identity, not insight, so usefulness has no sensible answer for them — asked anyway, the judge scored a correct brand name at relevance 65 and cut it. A wrong name is obvious to whoever reviews; a missing one blocks every generation downstream. - One bar becomes two: confidence 75, relevance 60. They fail differently. A claim that is not true is worthless at any relevance, but a claim that is true and merely an average instance of its field is still worth a row someone can edit. - The confidence bands no longer let one counter-example sink a rule that holds across the copy, unless the rule itself says "always" or "never". The anchors that were doing real work are kept, moved into the field purpose that owns them: a commercial policy without its numbers, a `dos` that is an adjective, an `avoid` that is another rule inverted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…led it Every competitor was rejected at confidence 25, seven in a row, with the same reason: no supporting evidence in the site's blocks. There was none to find. A brand almost never names a rival in its own copy, which is why competitors are researched on the web in the first place. The claims were labelled block-derived regardless, and the rubric tells the judge that a block-derived claim is judged against the site content — so it looked where it was told and correctly reported nothing there. Relevance scored 76-80 throughout: the judgement of the field was right, only the label lied. `preferFilled` already decides which source won, so it now returns that choice alongside the items and the specs are built from it. Splitting the value from its label is what let them drift; keeping them in one return is what stops it. The rubric also now says plainly that a research claim missing from the site's own pages is expected rather than a defect, with competitors named as the case. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A blog tool knows the brand's voice and nothing about its world. It cannot say what the store sells, at what price, what campaign is live, or what readers arrived looking for — so it writes around those facts or invents them. Meanwhile the site's virtual MCP already aggregates the connections someone attached to this CMS, and every one of them answers exactly those questions. The five generators now gather facts before writing: one tool-calling pass over the site's read-only tools, then the structured pass as before. It mirrors the two-phase shape `brand-research.ts` already uses for web search. Deliberately generic — nothing here knows what a catalog is. Tools are whatever the site has connected, described by their own MCP metadata, so a commerce API, an analytics property and a campaign system all arrive through the same door. Adding one later is attaching a connection, not editing this code. Two hard rules, because this runs with nobody able to approve a call: - READ-ONLY ONLY. A tool is offered only when it declares `readOnlyHint: true`. Unannotated tools are excluded: unknown risk is risk, and these sites commonly have a GitHub connection whose tools open pull requests. A generator that can open a PR is one that eventually will. - TOOL OUTPUT IS DATA. Results are third-party text on its way into published copy, so both the gathering prompt and the section it produces say so, and neither lets a result instruct. Pure enrichment throughout, like the research hook it mirrors: no virtual MCP, no read-only tools, a broken proxy or a failed call all yield "" and generation proceeds without it. A site with nothing connected pays nothing — no proxy, no model call, byte-identical behaviour to before. Bounded: 8 tool calls, 45s, 8k chars into the prompt. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
No blog tool declared `requiresAiBudget`, so an exhausted usage bar reported the
spend without stopping it — every one of these tools is real token spend, and
the grounding pass in the previous commit makes each click cost more.
Seven tools gated. The refusal is already wired end to end: the gate throws
`AiBudgetExhaustedError` ("… is paused: this organization has used its monthly
AI allowance"), the REST client carries the `code` through, and each blog
surface already toasts `err.message`.
Inert unless STUDIO_PLANS_ENABLED, like every other plan gate. Kept as its own
commit because it is a behaviour gate rather than part of the grounding, and is
worth being able to revert alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…xtract
Preencher read the catalog through `catalog-invoke`, which allows exactly two
VTEX loader resolveTypes. On a real site it answered 500 — the default listing
request sends `{ count: 24 }` with no `query`, `facets` or `collection`, and the
loader rejects it as unknown props. `fetchCatalogEvidence` swallows that, so the
extract quietly lost its best source of `keywords` and only a console 500 said
so.
Rather than fix the request, the path goes: the previous commit gave generators
`groundFromSite`, which asks the site's own virtual MCP the same question
through whatever connection answers it. The brand extract now uses it too, so
there is one way to reach a store's data instead of two, and it is the one that
does not name a vendor. A site with a VTEX Commerce APIs MCP, a Shopify one, or
an analytics property is served by the same code.
Removed: `blocks/catalog-evidence.ts` and its test, `EvidenceCatalogSchema` and
the catalog section of `renderEvidence`, `CATALOG_EVIDENCE_MAX_CHARS`, and the
`catalogChars` budget reserve in `selectBrandEvidence` that existed to hold room
for the sample.
Kept: the `catalog-invoke` route and its allowlist. The product picker invokes
those two loaders from four places and is unaffected — this removes one consumer
of that proxy, not the proxy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tions Authors and Categories were reference collections living inside the Context tab through a local master-detail panel. Context is about how the blog is written; those two are records the blog points at. Both editors are untouched and still reached from the content browser, which renders `RecordEditor` and `CategoryEditor` directly — this removes the second door, not the room. Automations takes their place as a locked tab: visible, never reachable, badged. The panel comes later. Falls out of the removal: `EntityPanel` had no other caller, nor did the `onOpenPost` / `onManageCategoryPosts` callbacks the two panels needed, and eight i18n keys in both locales went with them. `TabButton` grows a `disabled` state, since a tab that cannot be opened should not look clickable. `common.soon` is added at the position main already uses, so the pending sync reconciles rather than conflicts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…name A format's brief stored the bare `@Heading`, which names a component rather than a block. A site that overrides an app block has two of them — `blog/sections/blocks/Heading.tsx` and `site/sections/Blog/Post/Heading.tsx` — and the citation could not say which. `mentionableSections` papered over it by dropping one from the picker, so the override was uncitable, and `sectionResolveTypes` broke the tie by whichever sorted first. A citation is now the markdown link `[@Heading](<resolveType>)` — the shape `@decocms/shared/mentions` already uses for people, whose module gives the same reason: the label repeats and changes, so what it points at must not depend on it. It round-trips through the editor's existing link mark, renders legibly anywhere markdown is read, and the picker can stop hiding the override. A bare `@Name` still reads, because briefs written before this exist and the model still writes them: `citedSections` resolves it through the site's inventory, and `linkifyCitations` rewrites it on save, so one shape is persisted whoever authored the brief. A name the site has no block for is left bare on purpose — linking it would invent a target, and leaving it is what keeps `unknownCitations` able to report it. The model is deliberately not taught to emit resolveTypes. `sectionResolveTypes` exists so it never sees one and so can never invent one; it keeps citing `@Name` against a per-name inventory, and the web resolves that on the way in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Typing `@Heading` passes through `@H`, `@He`, `@Hea`, and the site has no block for any of them, so the "unknown citation" warning fired and cleared on every keystroke. The message appearing and vanishing under the editor — shifting the layout each time — reads as the editor lagging, which is how it was reported. Nothing was actually delayed: `rule.value` updates synchronously, and the 700ms autosave debounce only holds the network write. The warning was simply asking a settled-state question of text that was still being typed, so it now reads a value debounced by 600ms. Per row rather than per list: the open row's fields move into `RuleBody`, which owns the debounce. Shared across rows, the previous row's text would survive the window and warn about the wrong one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Suggest fired straight off the button with a fixed count of three and nothing the operator could say about what they wanted. The pillar panel already asks both, through a popover on its own trigger; formats now does the same. `guidance` joins the tool's input and the prompt, with the same standing the other suggestion tools give it: what the operator asked for outranks what the model read, since the existing posts and the brand profile are only evidence about what the blog is, and the operator is saying what it is for. Asked for a format this site has no sections to build, the brief says so rather than quietly proposing something else. `count` was already in the schema, capped 1-5, and was simply never sent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A pillar was defined as a durable territory — "Product updates", "Customer cases" — a lane the blog always returns to. That is not what pulls a post into existence. A moment does: Black Friday 2026, a line launching, a search the brand does not answer, stock that has to move. A campaign is a pillar with an end date. The durable half does not disappear from the product — it already lives in the brand context, in values, keywords and the commercial calendar. What was missing was the cut with a beginning and an end, and with something to point at. The object carries what a generated post will need: a status it moves through, an optional period, a closed `trigger` saying why it exists, an `intent` naming the objective and the targets it sells, and `guardrails`. Two of those names are promises about precedence, and the form says so under each field: `avoidComplements` adds to the brand's guardrails, `toneOverrides` replaces the brand's tone while the campaign runs. Targets are platform-agnostic: a URL identifies them, because it is the one thing every storefront has and the only one a reader can open. A product id would tie this to whoever issued it. The UI mirrors the Posts area, which already settled this pair — a board to move work through its states, a list to search and edit one at a time. Moving a card is much simpler than in Posts: a post changing status can cross the planning/live boundary and get its block key renamed, which is what `use-post-status-move.ts` exists for. A campaign is always planning, so it is one write. Only the in-flight set survives from that hook, so a card mid-write cannot take a second drop. Pillars are gone rather than deprecated, so `BLOG_PILLAR_SUGGEST` goes with them, as does the `pillar` slot in `BLOG_POST_DRAFT` and `BLOG_THEME_SUGGEST` — the latter was already dead, since both callers smuggled the pillar through `guidance`. Generation has no campaign slot yet on purpose: the wizard is the next step and will introduce one shaped for the new object, rather than inheriting a slot shaped for the old one. `RuleList` and `TermsInput` move out of `context.tsx` into `blocks/rule-list.tsx` — the campaign editor needs both, and duplicating them would mean maintaining the `@Name` citation parsing and the warning debounce in two places. Sites with blocks under `blog-manager/pillars/` keep them on disk with no reader. Nothing migrates them, by decision. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three problems, one shape: the form undid the choice, did not explain the
choice, or asked the wrong one.
The selector reverted. `useAutosave` re-seeds the draft whenever `initial`
changes identity, guarded only by `pending` and `isSaving`. The campaign editor
passed neither, so the decofile refetch that follows a save — carrying the
pre-rebuild block, because the dev server has not regenerated it yet — landed on
top of the click that caused the save. Pass `isSaving` the way post, category
and record editors already do, and give the absent block a stable empty seed
instead of a fresh `{}` per render.
Picking from a closed set now saves at once rather than after 700ms. There is no
next keystroke to wait for, and the wait was the window the stale echo arrived
in.
Objective and trigger carry a help popover. They are the two fields generation
leans on hardest and their labels are jargon; the legend sits beside the chips
rather than under them, where twelve lines would bury the form.
Targets no longer offer "product". Choosing between a product and a category was
a false choice: whatever the target is, the campaign still needs the products
the copy may name. So a target is a category or a collection — the slice being
argued for — and highlighted products are their own list beside it, always
asked. Both carry name, url, id and description, and both are typable: the store
pickers are a shortcut, never a gate.
Categories come from the store's tree, extracted out of the product picker so
both surfaces share the fetch and the filter. Collections stay manual — nothing
here lists them. Products come from the picker, which now also parses the main
category and a short description out of a payload that already carried them.
Products are copied into the block rather than referenced by id: generation
reads them months later, and an id only answers while that storefront is up.
Fixes the picker's empty browse, which answered 500 `Unknown props: {"count":24}`
— four of the five request branches pair `count` with a discriminator and work;
the empty one sent `count` alone, matching no variant of the loader's props.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Targets and products were columns of open forms — one entry filled the viewport, and adding a second was something you had to believe in rather than see. They are records with one obvious title and several fields behind it, which is what the brand's rules already are, so they now read the same way: a row per entry, click to open, one open at a time. That shell is extracted out of RuleList rather than copied. Only the shell — each list owns which row is open, because they differ in what deleting one should close. Adding is now section-level and doubled: browse the store, or add by hand. The category picker used to live inside a row, which meant the first target had to exist before the store could fill it, and made a second target hard to find. A category target keeps the URL the store's own tree reports. It was being rebuilt from the path, which dropped the origin and stored a pathname where a reader needs a link. A product carries up to three images instead of one. The payload always had them; the picker took the first and dropped the rest, so changing the image meant going back to the store. The cap lives with the model — the picker reports what the store has and the campaign decides how much to keep. The trigger and objective legends open in a modal. As a popover they rendered as a cramped strip over the field they were explaining. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Add image" did nothing visible. The row was created and destroyed in the same render: the editor re-reads the draft through `scanCampaigns` on every pass, and `readProductImages` dropped empty strings, so the empty slot the button had just appended was gone before anyone could type into it. Blanks are kept now, with a test pinning it. The objective legend was five one-liners, which is not enough to choose with — "awareness" and "education" read as two words for the same shrug. Each objective now carries the four things that actually separate them: what the post is about, how the product shows up in it, how hard the CTA pushes, and what you would measure. Education and conversion also say which trigger they usually pair with. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The image slot was a bare URL input — the one place in the product that could not reach images the brand had already uploaded. It is now the Studio image field, the same one the post cover, the SEO image and the author avatar use: browse the org's buckets, paste an address, or drop a file. Compact, because three of these sit inside one collapsed product row. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n systems Filling a campaign in by hand means knowing what sells, what has stalled and what people search for and do not find — all of it in the brand's systems rather than in anyone's head. So a campaign now starts from a seed: some keywords and a sentence about the moment. BLOG_CAMPAIGN_SUGGEST reaches the site's MCP connections from that seed. The grounding pass already existed for the brand extract and its prompt was already written for this — it asks what is selling, what readers arrive looking for, what promotion is in flight. It gets a deeper budget here, because sweeping analytics and a catalogue is not the same as reading one site. It proposes up to three candidates and the person picks. Up to: the model is told that a seed pinning the campaign down deserves one answer, not three near-duplicates the person has to read to discover are the same. A separate pass reviews each candidate — verdict, rationale, risks. Separate because the model that just wrote the campaign is the worst judge of it. A failed review keeps everything as workable rather than discarding work someone is waiting on, the same rule the claim judge follows. Nothing invented. The prompt forbids a target or product the tools did not report, and `groundedOnly` enforces it after the fact: with no store data every proposed target and product is dropped and the dialog says so. A made-up product URL is worse than a blank field — a blank field asks to be filled. Seeds are blocks, not a form that vanishes. The answer changes as the store does, so a seed worth writing once is worth running again next quarter; campaigns point back through `seedKey`. The closed sets move to @decocms/shared, because both ends now validate against them and restating them server-side would mean a trigger added to one end and silently rejected by the other. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A grounding pass had 45s; a single MCP tool call had the shared default of 120s. The inner budget was nearly three times the outer one, so one slow connection ran out the clock and the abort discarded every finding the other connections had already returned. Three things were wrong, and they compounded: The per-call timeout now takes a share of the pass it runs inside. A call that cannot answer in a fifth of the budget has already cost more than it is worth. A failed tool is now an answer, not the end. `toolsFromMCP` rethrows, which is right for chat — a person sees the error and retries. Nobody is watching this one, and the system prompt already tells the model to note a broken tool and carry on; it just never got the chance, because the throw unwound the whole call. It now reads an error result and keeps going with the connections that do answer. The pass budget goes to 75s. Grounding against a real store with analytics and a catalogue is slower than the single site the original number was measured against. The log line was a thirty-line DOMException dump, every one of its twenty-five error constants included. It is one line now, and says plainly when the cause was the pass giving up rather than something breaking. Extraction itself was never affected: grounding is best-effort by construction and returns "" on any failure, so the brand still extracted from blocks and research. What was lost was the store data, silently. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The timeout was ending the pass the worst possible way. The store answered, the abort fired, and every finding went in the bin — then the dialog told the person no connected system had answered, which would send them off debugging a connection that works. The loop now stops gathering at 60% of the budget and, if that left it mid-tool-call with no prose written, takes one tool-free turn over the same transcript to write up what came back. The abort stays as a safety net rather than a schedule. With the work no longer thrown away, the campaign budget goes to 120s. The cost of a bigger number is now someone waiting, not someone getting nothing. `ran: boolean` becomes a five-value outcome, because the gap message was collapsing things that ask for different responses: a site with no connection wants one, a connection exposing nothing read-only is a permissions question, and a pass that timed out was being answered. A timeout that got results now says how many came back instead of claiming none did. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… one A real run told us what the design could not: 24 calls over 314 tools, 112s of a 120s budget, and every single call went to the catalogue. The analytics connection the operator had attached was never consulted — which defeats half the point, since what readers search for does not live in a catalogue. Four of those calls were the same product search repeated. The pass now says how many separate systems its tools come from, grouped by the connection they arrived through — the names are prefixed by connection id, not by vendor, so that is the only reliable grouping. The model is told to spend across them, and that finishing without touching one is an incomplete answer. It is also told not to repeat a call, and that it has far more tools than it has time. Three hundred tools is not a menu to work through. Latencies were never the problem: every call came back in one to five seconds, nowhere near the per-call ceiling. The budget went on wandering, not waiting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first real campaign came back with targets that had an id and a name and no URL, and no products at all. That was not the model failing — it was the model obeying. VTEX's catalogue tools return category ids and names, never a storefront address, and the rule is never invent the store. With no origin known, there was no honest link to give. The product schema asks for a URL and images, neither came back, so it returned an empty list rather than a made-up record. Its own trigger note names the SKU it could not link. The same gap leaks the other way: the picker invokes loaders through `catalog-invoke`, whose origin is the deco runtime, so the URLs it hands back carry the VTEX account's internal host. That address works, which is what makes it dangerous — nothing fails, it just gets published. Nothing in Studio held the store's public domain. `previewServerUrl` is deliberately the runtime host, and the one place that knows the production domain reads it during a deco.cx import and throws it away. So the brand context gains `storeUrl`, validated with the `sanitizeSiteUrl` that already exists in shared, filled by the brand extract like any other field and scored by the same judge — with a purpose that names the internal hosts as wrong answers however confidently a site reports them. `reHome` puts a picked URL back on that domain: keep the path the store authored, drop the origin we accidentally asked through. Applied where the host leaks — campaign products, category targets, and the rich-text product link, which is the one that ends up as a published href. A blank store address leaves every URL untouched; a brand field nobody filled must not make a working link worse. Generation learns the distinction that was missing: joining the store's address to a slug a tool reported is building a link, not inventing one. Guessing the slug is still forbidden, and a host from a tool is never usable. It is also told that a product with an id and a name is worth returning without a URL — the editor completes it — which is the direct cause of that empty list. Separately: `contextForTools` was dropping `specialDates` and `competitors` while three prompts rendered them, so a seasonal campaign was flying blind over the brand's own calendar. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Deleting one meant opening it first and finding the button in the editor footer, which is not where anyone looks for it. The trash now sits on the card and on the list row, revealed on hover, the way the post board's does. It asks first, which the post board does not. There the delete is soft — the post lands in Archived and nothing is lost, so a confirmation would only be noise. A campaign has no such lane: the block goes, and the targets, products and guardrails inside it go with it. The dialog says that, and says what survives, because the thing someone actually fears is losing the posts. All three entry points route through the same confirmation, so the editor's own button no longer deletes on a single click either. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three things the first grounded campaign got wrong. A collection target was asking for a URL and warning when it had none. It has none because there is nothing to have: a collection is a cluster the store filters by, not a page. The id is the identifier. The URL field now hides for collections, the warning goes with it, switching a target to collection clears any URL carried over from when it was a category, and the server blanks it regardless of what the model returns — the prompt asks, this makes it true. It suggested one product. One product reads like an ad for that product; the field is a shelf. The model is now told to list every product the tools surfaced that fits, and to say so in the trigger note when the catalogue really only offered one, rather than leaving a thin answer looking like a choice. Images came back empty. The ask for them was buried in a sentence about products generally, so it is now per-product and explicit. It also says why they must be copied character for character: an image lives on a CDN whose host has nothing to do with the store's address, so the composition rule that fixes product links would break every image it touched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A campaign came back with one product, no image and no link, and two targets reported as "dropped: nothing in the store data backed them" — in English, inside a Portuguese interface. Two faults, and the interesting one is not the language. Grounding was prose only. The raw tool results were summarised to 8k chars of markdown and thrown away, and a summariser has no budget for three 150-char CDN URLs per product, so the images never survived the write-up. Nor did the slugs — and without a slug there is nothing for "building a link is not inventing one" to join, so `url` came back empty. Correctly. Worse, the oversized results — the catalogue's richest — were truncated to a 400-token preview pointing at `read_tool_output`, a tool this pass does not offer, with the full JSON parked in a map that was discarded. And `groundedOnly` substring-matched proposals against that summary, so a category the summariser paraphrased was dropped as ungrounded. The guarantee was being checked against the paraphrase rather than against the answer. So keep the raw results, transcribe a catalogue out of them on the cheap tier (copy or omit, never compose), and hand that to the generator as a shelf to choose from by id. The slug becomes an address in code, via `reHome` — now shared, since both ends need it. Grounding checks ids against the catalogue and falls back to the raw bytes, never to the prose. The gaps become codes with counts; the UI words them from the locale the person set, and `modelSummary` carries the English for an agent calling the tool. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The extraction pass came back with 68 categories and zod threw the whole object away, three times, because the schema said `.max(60)`. A store with more categories than a number I guessed is not a malformed response — the cap belonged after the parse, not in it. Trim in code; the schema only describes. The same run exposed the window: 1.3M chars of tool output, of which the pass read 120k. This store lists its categories before it lists what is in them, so every product sat past the end and `products` came back empty — a real catalogue, correctly transcribed, with the half that mattered out of frame. Read the evidence in parallel chunks instead, split on the entry boundary so a JSON record is never cut in half, and merge by id. `slug` stops pretending to be a slug. A catalogue API answers with a whole URL on an internal host, and asking the transcriber to strip it was asking it to edit — which is the one thing this pass must not do. It copies; `reHome` launders the origin afterwards, which it was already doing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Asking a model to copy a CDN URL across from a catalogue is asking for an image that does not load, and a `.max(3)` on the field it copies into is asking for the whole campaign to be rejected when it copies four. Both were there. So it stops copying. It returns the `id` and the argument for why the product belongs; the address and the images are attached afterwards from the catalogue itself, in code. The id already identifies the record — retyping the rest was a transcription step that could only lose. That also pays for the shelf. The catalogue goes into the prompt without its image URLs, which is most of its bulk, so eighty entries fit where a dozen did, and it states its own size: returning one of forty-seven is visibly a failure to choose rather than a plausible answer. An ungrounded proposal is untouched — a name and a description with no catalogue behind them are still worth having, and that is the no-connection path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # apps/api/src/tools/blog/brand-extract.ts
`images` kept a `.max(8)`. A product with nine images rejected its whole chunk, three of four chunks died that way, and the catalogue came back with zero products — the identical failure to the one fixed two commits ago, left in place because that fix went looking for the cap it had already found instead of for the pattern. The rule is now written where the next schema gets added: on model output a length cap is a rejection, so nothing here carries one. Caps on what is KEPT go after the parse; caps that reject go on `inputSchema`, where the sender can be told. The same run showed the other half. Forty-nine tool calls leave a transcript too long to summarise, so the write-up came back empty — and an empty write-up printed "No connected system answered. Leave `targets` and `products` empty", directly above a section holding the store's own catalogue. The prompt contradicted itself, the result reported `grounded: false`, and the gap banner said nothing useful was found. Prose is no longer the only thing a pass returns, so none of those three read it that way any more. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The chunks see different halves of the same product. A catalogue listing names it and carries no image; the SKU files that carry the images arrive tens of thousands of characters later, in another chunk. `mergeById` let the first sighting win whole, so every product listed before it was photographed had its images fetched, transcribed, and then thrown away at the merge. Merge field by field instead: a filled field is never overwritten — the earliest chunk read it nearest its source — but a blank one stays open to a later answer. Also raise the chunk budget to cover the evidence budget. Four passes read 480k of the 600k kept, so an eighth of what the store said was gathered, stored, and never looked at. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lly typed Generation told the writer `- Callout — Highlighted note, tip or warning` and nothing else. The props came from a seven-kind schema hardcoded in the tool, duplicated in a second list on the web, and written out by a seven-case switch that encoded each block's storage convention by hand. That was only ever right for the blocks the blog app ships: a site with its own `Callout` got its resolveType matched by name and its props filled from assumptions, which saves without error and renders empty. The site's real schema was there the whole time, in `meta.schema`, resolvable, and untouched by anything in the generation path. Now the format's `@Name` citations pick the blocks — `citedSections` already parsed them and nothing ever used the result — each block carries its own schema into the prompt, and what comes back is validated against it. The switch is a spread. Campaigns were a dead end: `toneOverrides` and `avoidComplements` were written, stored, and read by nothing. A post is now written FOR one, and takes from it the moment, the objective, the products it may name and the links and image addresses to use for them. The tone it overrides is the tone the post is written in. The brand context it travels with is half the size. Four of the sixteen fields were already being dropped unread by this prompt; the rest competed with the brief for attention. What is left is what changes how a sentence comes out. And a draft is born with a cover. Not via `generateImageCore`, which persists to object storage served behind a session check — fine for a chat attachment, useless for a published post, where the first anonymous reader gets redirected to the login page. It goes where a cover uploaded by hand already goes: the organization's own bucket, whose public URL is a real CDN address. A block prop typed `image-uri` can ask for one too, by sentinel, capped at two a post. Ideas go, as asked — the tray, the lane and the step — and `BLOG_THEME_SUGGEST` with them, since nothing else called it. Ideas already saved are left in the decofile rather than deleted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The brand context tab, the campaigns inside it, post generation and link suggestions now appear only for an organization that turned them on. Off by default, next to the existing blog-blocks toggle in Settings → General. The flag is one line in `OrgFlagsSchema`, which is what that file's header promises: storage, both settings tools and the web hook derive from it, so there is no API change to make. The two blog toggles share one "Blog" section rather than taking a heading each — they are both answers to how this org blogs. `activeCollection` is local state and the context tab is reached by clicking rather than by link, so nothing can navigate into a hidden tab. What could happen is someone standing on it when an admin flips the flag. Normalising the collection once at the source handles that, and spares the two readers downstream a guard of their own. The flag governs what is shown, not what is callable: the tools stay reachable over MCP. A feature flag gates product behavior, not access. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every section of a generated post was dropped — 23 of 23, post empty. Not the
model: the schema it was validated against demanded a property the model is
deliberately never shown.
deco wraps a block's config as `{ properties: {__resolveType}, required:
["__resolveType"], allOf: [{$ref: Props}] }`, and `resolveSchema` strips every
`__`-prefixed key from `properties` while carrying `required` through whole.
What came out the other end required `__resolveType` and declared no property
for it, so ajv rejected any object that did not carry it — and the writer could
not have known to write it, since keeping the resolveType away from the model is
the point. The caller stamps it afterwards.
So `required` is now filtered to names the pruned schema still declares.
The reason this took a guess to find is the second fix: a dropped section said
nothing about why. It now carries its reason — not a block this site has, not
JSON, not an object, or the ajv message — tallied by reason so the log groups
instead of flooding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e blocks Two symptoms from one real run, and only one of them is diagnosable from here. The product shelves and cards came back full of slugs that render nothing. Those props carry `"format": "dynamic-options"` — a loader resolves them against the live catalogue, so the value is an id the store owns and the writer has no way to know it. Nothing said so, and the pruned schema had dropped the `options` loader path, leaving the field looking like a free string. The path now travels, and the prompt says plainly that an empty shelf is one someone fills in two clicks while a shelf of composed slugs renders blank and looks filled. The list came back as an object where the block stores a newline-joined string. That one should have been caught: a probe through the real pipeline rejects it with "data/items must be string", so the schema that reached the writer was not the one the site publishes. Which means the thing this tool cannot see is exactly what it needs — so it now logs the props each block offered, name, type, enum and widget. Two rounds of "why did that section come out wrong" went to guessing for want of that line. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The log answered the open question: `List: items=object`. Not the model guessing and not ajv failing — the schema the site publishes genuinely says object, so a map was the obedient answer. It came through unpruned, which means that object declares no properties: it is `resolveSchema`'s last-resort branch, the one whose own comment says "ref points to a def with no type we recognize, treat as object". A guess, and the writer acted on it. So each block now also carries one of itself as the site already stores it, read from a live post. A stored section is what actually renders, which makes it the authority a derived schema only approximates — and the prompt says so: where the two disagree, follow the example. Live posts only. A planning post may have been generated wrong, and copying our own mistake back in as the example would make it permanent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Last run the writer did what it was told — the `List` came back as the newline-joined string the site stores — and then I dropped the section for it: `data/items must be object`. Telling it to follow the example and validating against a schema that contradicts the example is a contradiction that costs the section either way round. Both halves are evidence, and where they disagree the stored block is the one that renders. So a property the example disproves loses its `type` and `enum` before the schema reaches either the prompt or ajv. Narrow on purpose: the example disproves what it covers and nothing else, and a stored `null` disproves nothing — it is an empty slot, not a type. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… has Read the site's own stored blocks instead of reasoning about them, and the answer was sitting there: a product field holds a bare id — `"1948858"`, and a shelf holds a list of them. Not the `vtex:sku:` form its own doc comment advertises. The loader behind it falls through to a fulltext lookup for anything unstructured, which is why plain ids work at all. The campaign cannot supply that id: what the catalogue reported is a stock code, not what the storefront indexes by. What the campaign does have is the product's address, and the slug in it goes down the same fulltext path. So each product now carries a `reference` derived from its URL — in code, because composing a slug is the kind of thing a model does plausibly and wrongly, and this one is a substring. The rule it replaces said to leave these fields empty. That was the safe answer while the shape was unknown; it is the wrong one now that a usable handle exists. What stays is the prohibition that matters: never take the value out of a block's example. That example points at a real product — a different one — and a card showing the wrong product beats a card showing none only in looking deliberate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…'s view
`resolveSchema` is the form renderer's view of a schema. It exists to decide
which field to draw, so it flattens what a field cannot show — and a prop
declared `string | string[]` comes out of it typed `object`, because the
last-resort branch treats an unrecognised union that way. The writer was handed
that and wrote an object, which is how a list the site stores as one
newline-joined string came back as a map, twice.
The meta had the real schema the whole time. deco wraps a block as
`{ allOf: [{$ref: Props}], properties: {__resolveType}, required: [...] }`, so
the props are one hop in, and that hop keeps everything the writer needs and
none of what the form needed: the `anyOf`, the `format`, the `options` loader,
and the descriptions — including the one that spells out the reference format a
product field takes.
`$ref` is followed through the meta's own definitions, bounded by depth and by
size, because a product ref expands without end. `__resolveType` is dropped on
the way out: the caller stamps it, and a writer that never sees one cannot
invent one.
The adapter that flattened the editor's view goes with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # apps/api/src/tools/index.ts # packages/shared/src/tools/registry-metadata.ts # packages/shared/src/tools/tool-io.ts
`@format` in a deco schema picks an editor control — `textarea`, `dynamic-options`, `image-uri` — not a value constraint. ajv has never heard of any of them, ignores them, and warns once per compile: a line per block per run, and a red herring every time someone reads the log looking for why a section came out wrong. It has cost two readings already. Only the copy ajv sees loses them, cached per schema so the validator's identity cache still hits. The prompt keeps `format`, because it is how the writer learns a field is loader-driven, and the image scan reads it to find where an image was asked for. The type under the hint is still enforced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…read A site whose section resolves the product reference in its own loader stores the id and nothing else. `product-loader-utils` knows three shapes and all three carry a `__resolveType`, so against a bare id it answered `[]` — a correctly filled block showed up blank in the editor — and on write it wrapped the id in a VTEX loader ref, handing the site a shape its own component cannot read. The file's own header promises to "preserve the stored shape on write"; for this shape it did the opposite. That broke the block through the editor, with no generation involved. With the fourth shape understood, the same helper can place generated products too. So the writer stops composing a product slot: it chooses ids from the campaign and the slot is written in whatever shape the site already uses, read from one of its own blocks. Where no block shows how this site stores them, the slot is left alone — guessing a loader ref is how this bug happened once. The id itself now comes from the campaign and has to be the right one: the catalogue extraction spells out that it wants the id the storefront indexes by, not a reference code or a stock code, because those resolve to nothing. The slug-from-URL bridge goes — it stood in for an id that was never captured. And a table of canonical examples for the blocks Studio draws its own editor for. That editor is the contract, and several of those shapes are ones no schema states plainly — a list keeps its items as one newline-joined string, and six blocks keep theirs as JSON encoded inside a string. It is the floor only: a block from the site's own published post still wins, because two sites legitimately store the same block differently. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A prop declared `string | string[]` is valid either way, so the schema cannot decide it and the writer picks one — as often as not the one the schema lists first. For a list that this site keeps as a single newline-joined string, an array is accepted, stored, and wrong. The site already decided: whatever a block it renders holds is what its section reads back. So where the two forms are the same list of strings and the writer chose the other one, it is converted — joined or split — before validation, in code, with no model involved. Only a list of strings, only where the two disagree, and only when a block of the site's is there to go by. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What is this contribution about?
The blog's editorial context was one JSON block and one AI pass. Two things were tangled in it: who the brand is, which holds whether or not it ever publishes a post, and how the blog is written, which belongs to the blog and evolves with it.
The split
blog-manager-brandblog-manager-context(new)Two chained tools:
BLOG_CONTEXT_EXTRACTtakesBLOG_BRAND_EXTRACT's answer as input, because writing rules read out of a site's copy are only as good as the understanding of whose copy it is. A failed brand pass aborts rather than inferring rules against a blank profile.No migration. Sites written before the split keep working: the writing rules fall back to the brand block while
blog-manager-contextis absent, and each save writes only the fields its block owns — so the old block sheds the moved fields on its next write. Same tolerancenormalizeBrandRulesalready gives the legacystring[]rule lists.Six new fields
specialDates,keywords,differentiators,commercialPolicieson the brand;vocabularyandvoiceExampleson the generation side. All render through the existingRuleListexcept example phrases, which get a sounds / doesn't sound toggle — a counter-example is what pins down how far the imitable ones go, andpost-draftrenders the two sides as separate lists so the model doesn't imitate both.Three new evidence sources, same one button
seo-evidence.ts) — never read on purpose before. The site-level default lives onsite/apps/site.ts, which is neither a post nor a page, so no evidence tier reached it; page SEO leaked in unlabelled and undeduped. It is a site's most deliberately written copy. Deduped on rendered content, because a templated PDP title repeats byte-for-byte.catalog-evidence.ts) — through the existingcatalog-invokeproxy and the product picker's own request builders. The VTEX account and credentials stay in the running site's app; nothing is fetched cross-origin and no new allowlist is introduced. Gated on the manifest, so a non-VTEX store costs no round trip.brand-research.ts) — broadened past competitors to dates and differentiators, the three things a site structurally cannot say about itself. One search call plus one structuring call. Citations are captured (researchSources) so a human can check an invented date.Pages now rank home → institutional → commerce. A PDP is a template with a product name substituted in, so the thousandth teaches nothing the first did not.
Also
soon-overlay.tsxdeleted; it had no other caller).How did you verify your code works?
bun run verifygreen: fmt, oxlint, knip, 9613 tests, 0 failures.tsc --noEmitclean onapps/webandapps/api. Tool contracts regenerated.41 tests added or rewritten, all in the paths where this change can lose data or silently drop evidence:
blog-data.test.ts—splitBlogContext(legacy block seeds the rules; the context block wins once it exists; the brand half never carries a writing-rule field);pickBlogFields(the save allowlist that makes the legacy block shed moved fields);applyExtractResult(empty vs replace, whitespace counts as empty, a blank editor row is not an answer, replace keeps an unanswered field);normalizeVoiceExamples(new shape, the{name,value}shape this field briefly had, bare strings,soundsdefaulting); the sampler's new page ordering and its budget reservation.seo-evidence.test.ts(new) — lazy-wrapped SEO,seo: null, site-nested vs standalone, PDP/PLP role detection, content dedupe, the two-commerce-page cap.catalog-evidence.test.ts(new) — both loaders required before any round trip, schema.org shape variants, missing price/brand dropping from the line rather than the product.One
verifyrun showed 1 failure that did not reproduce on the next three runs and that I could not capture. It was not in the blog paths; if it shows up in CI it is worth a look, but I could not tie it to this diff.Screenshots/Demonstration
None — and this PR has real UI changes, so that is a gap, not an omission by choice. I never ran the app: six new editable sections, the fill-mode dialog and the phrase toggle have only been exercised by type checks and unit tests. Someone should run the steps below and attach evidence before this merges.
How to Test
devhybrid, open a site with posts → Contexto → aba Marca.blog-manager-brandon disk): the tabs should open already populated, with tone/dos/guardrails read out of the old block.blog-manager-brand.jsonandblog-manager-context.jsonexist, each holding only its own fields.catalog-invokecall fired on a VTEX site — then repeat on a non-VTEX site and confirm the extract still completes with no catalog.Migration Notes
None. The split is read-tolerant by design (see above) and no block needs rewriting.
Review Checklist
Two things deliberately left out, both worth their own PR:
requiresAiBudget(define-tool.ts:100), so an exhausted AI bar does not stop Preencher — and this PR makes each click cost more. It is a behavioural gate, unrelated to the fields, and you want it revertible on its own.categoriesis still generated by the model and discarded by the UI, exactly as before this PR.🤖 Generated with Claude Code
Summary by cubic
The blog's AI tools are rebuilt: the single brand/SERP context splits into a brand block and a separate writing block, every generator grounds itself in the site's own connected MCP systems, content pillars become campaigns, and drafts are written against the site's real block schemas.
type/enum, required props are filtered to the pruned schema, editor-onlyformathints are stripped from the copy ajv validates, astring | string[]prop is joined or split to match the form the site's own stored block uses, and a dropped section says why.storeUrl; the catalogue is read in parallel chunks, merged field by field, and a canonical shape table covers the blocks Studio edits itself.image-uriprop can ask for one by sentinel, capped at two a post; the other props a loader owns (dynamic-options) are flagged so the writer leaves them empty.blog_ai_enabled), off by default, gating what is shown rather than what stays callable over MCP; blog generation stops on a spent AI allowance.[@Heading](<resolveType>)links; bare@Namestill reads and is rewritten on save, and the unknown-citation warning waits 600ms.inputSchema; caps on what is kept run after the parse.Written for commit 9a22ab6. Summary will update on new commits.