Agentic Marketing Team · Contract Handoffs
13 Agents · 3 Departments · Rejections With Teeth
Agentic Marketing Team: Three AI Departments With Enforced Handoffs
A brief goes in; positioning, campaign, press kit, and a designer-ready brief come out. Three departments hand work to each other under an enforced contract — every claim carries its source, adversaries can send work back, and the console shows every dirham and every rejection as it happens.
Personal Project · Working build
13
Agents, Three Departments
Marketing, creative, and comms run in sequence: 4, 4, and 5 agents, each department ending in an adversary with real power
$1.36
Cost Per Run, Verified
The demo run split by department: $0.20 marketing, $0.49 creative, $0.67 comms at current API pricing, retries included
18:33
Brief to Package
Recorded wall clock on the demo run. Its capture replays the full event stream, at real recorded cost, in about 35 seconds
2
Rejects Caught in One Run
A live run with planted defects: the skeptic killed an unverified superiority claim, the critic killed an unverifiable photo dependency
The Sibling System
Marketing Debate is a panel: seven agents argue a brief in parallel and produce a decision, with every inferred number marked in the document. This is the org that does the work: departments in sequence, deliverables instead of a verdict, and the same provenance discipline promoted from a marking in one document to a contract that survives crossing team boundaries — with adversaries empowered to send the work back.
The Handoff Contract
What actually travels between departments
The point of the system is the moment work leaves one department and arrives at the next, carrying its assumptions with it. The contract is what makes that moment trustworthy: an unverified figure cannot silently become a fact by being handed to someone new.
Every artifact travels with its claims
A department hands forward a summary, its artifacts, a handoff note, and a claims register. Each claim is tagged at the source: stated in the brief, or inferred — and an inferred claim without a basis string fails validation before it can cross the boundary.
Registers are inherited verbatim
Downstream departments receive the upstream register word for word — never restated, never renumbered. A figure the marketing skeptic flagged as unverified is still flagged when the comms drafter writes the press release, three handoffs later.
Zod at every boundary
Nothing unvalidated crosses a department line. Agent replies are schema-checked with one corrective retry; the assembled department output is validated again before the next department reads it.
Qualifier text is claim-bearing
Disclaimers and legal-safe lines are validated against the cleared-language register exactly like headlines — because a live run proved a qualifier is where an unverified process claim sneaks through. Where a permitted substitute exists, it is used verbatim.
The Org
Thirteen agents, each department closed by an adversary
Sonnet 4.6 does the building. The two in-department adversaries run on GPT-4o — a different provider on purpose, so the work is judged by a model that did not write it. The press-kit drafter runs on Opus 4.8.
Marketing
Audience → Positioning → Media → Challenge
The skeptic closes the department: it names the single load-bearing assumption, the collapse scenario if it is wrong, and the cheapest verification — then writes the handoff creative works from.
Creative
Concept → Copy → Art → Critique
The critic runs one test on everything: could this ship unchanged for a competitor? It also catches uncleared figures sneaking into headlines — quoting the sentence is mandatory, an accusation without the quote is not a finding.
Comms
Intake → Strategy → Draft → Review → Design Brief
The reviewer can return the package to creative once per run. The design-brief writer runs after the verdict settles and turns the approved package into a paste-ready prompt a designer can execute without reading anything else.
Failure and Recovery
Work flows backward when it deserves to
An org that can only pass work forward is a conveyor belt. The interesting part is rejection: an adversary that can send an artifact back with a quotable reason, and a reviewer that can send a whole package back a department — both on a budget the engine enforces.
Reject: one retry, in-department
The skeptic and critic return a verdict with a target and a quotable reason. On reject, the targeted agent reruns once with the rejection appended, then the adversary re-verdicts. On the console it holds on screen: a coral curve from adversary to target, a REJECT · RETRY 1/1 tag, and the reason at projector scale until the department completes.
Bounce: the package goes back
The comms reviewer can return the work to creative once per run. Concept and copy rerun against the objection; art direction and critique are retained; and the design-brief writer is deferred until the revised creative lands, so the final deliverable is never built from artifacts the reviewer just sent back.
Bounded by the engine, not the prompt
One retry per agent per run. One bounce per run. A second rejection is logged and the run continues. The prompts request; the engine grants — a misbehaving reviewer cannot burn the budget, because the bounds live in engine state it cannot touch.
From a Real Run
A superiority claim meets the pipeline
A live run on a realistic launch brief for a 180-unit Abu Dhabi development — with a claim the client had only confirmed internally. Everything below is the actual output of that run, unedited, following one claim from the brief to the final sign-off.
The client brief
“Halcyon Quarter is the most energy-efficient residential development in Abu Dhabi.”
The positioning direction arrived as a superiority claim, confirmed only by the client’s own internal modelling — exactly the kind of number that usually hardens into a headline.
Positioning Lead — marketing
The claim is tagged GAP with a named verification path, not repeated.
“That modelling is the starting point, not the finish line… it must be independently verified before it appears in any public-facing material. Commission an independent assessment — Estidama Pearl Rating System is the Abu Dhabi standard.”
The Critic — creative, reject fired
A headline died for inventing a checkable statistic.
“Headline fails swap test with factual challenge: ‘Al Reem Island has 30 residential towers. One of them was engineered to run cheaper than the rest.’” The copywriter had invented the tower count; the rewrite came back clean.
Cleared-language register — comms
The superlative is quarantined to permitted substitutes.
“Energy-efficiency claim — direction only, no superlative. Permitted: ‘Engineered to cost less to own’ or ‘designed to reduce annual operating costs.’” Internal figures — lead capacity, the 90-day target — are marked internal use only, not for press.
The Reviewer — comms
APPROVED WITH CHANGES.
The package ships only if the embargoed figures stay embargoed and the certification is confirmed before launch — with the substantiation register as the legal review list, naming who signs off on every held claim.
The Bug the System Caught in Itself
The disclaimer was the leak
The most instructive failure was not a headline — it was legal-safe small print. The fine text under every layout asserted a certification process nobody had started, and the reviewer waved it through because it looked like protection.
Found
In one run the reviewer quoted a banned phrase as protection: “it is qualified by the legal-safe qualifier line… independent third-party certification is pending” — while its own register forbade exactly those words. The copywriter had invented the pending certification; the reviewer treated the qualifier as a mitigant without checking its wording; the design-brief writer inherited it into the mandatory disclaimer of every layout.
Fixed
Qualifier, disclaimer, and legal-safe text became claim-bearing at two levels: the design-brief writer must validate every qualifier against the cleared-language register and may never assert a process or timeline without a substantiation entry, and the reviewer’s check explicitly covers qualifier text — an unverified process assertion is a must-fix, and a qualifier is not a mitigant if its own wording is uncleared.
Proved
A comms-only rerun against the same flawed creative input flipped the verdict: NOT READY — BOUNCE TO CREATIVE, quoting both leaks. The corrected disclaimer states the truth instead: “Independent third-party certification has not yet been initiated.”
The Console
Watching an org work, not watching a spinner
The dashboard is built for a projector at the back of a meeting room: department rings that read from a distance, a persistent mark where a rejection happened, and the money on screen next to the work it paid for.
One stream, every run
A single server-sent event stream carries every run to the console; a heartbeat doubles as the liveness signal, so the connection indicator shows connected, reconnecting, or down — and recovers through a server restart without a reload.
Cost and time, per department
Token counts accumulate per model call, retries included, and price against a rates file. The console shows spend per department and per run, elapsed timers per department, and a stats grid where “unsupported claims” is derived by scanning the actual register — not hardcoded.
Recorded time, not playback time
Replays show the original wall clock — the real 18 minutes, labelled “recorded” — never the accelerated 35 seconds. The demo is honest about being a replay.
Stop, with a confirm
A running live run can be stopped: two-step confirm, the in-flight model call aborted server-side, and the run lands in a distinct stopped state with completed departments, their artifacts, and their costs retained — still replayable up to the stop point.
Every live run becomes a fixture
The capture file is written event by event as the run executes, so even a run that dies halfway is a replayable fixture. The demo replays are captured real runs, not scripted animations.
The close is a deliverable
On completion the canvas settles into a summary — time, cost, artifacts, claims, and rejects caught given equal weight — with two downloads: the full package as markdown, and the design-brief prompt ready to paste into a layout tool.
Screenshots
The console, live
Screenshot
Run Console: Live Replay

Run Console: Live Replay
The console mid-run: three department rings with agent nodes, per-department cost and timers below the canvas, and the event log streaming every transition.
Screenshot
Completion Summary

Completion Summary
The close: recorded time, cost, artifacts, claims on the register, and rejects caught at equal weight — with the full package and the design-brief prompt as downloads.
Decisions the Measurements Forced
Four places where a live run changed the design
Each of these has the same shape: a real run failed or measured badly, and the failure was traced to a design choice rather than a model mood.
Prose left JSON after two dead runs
Agents originally returned artifacts as markdown inside a JSON string. Two live runs died on the same agent — a 3,400-token media plan cannot be escaped reliably — so the contract changed: a small JSON head, then the body below an ===ARTIFACT=== marker, merged before validation. The same agent that failed four times passed first try on the new format.
Timeouts came from the stopwatch
No timeout constant was set until runs were measured: max single call 95 seconds, later 109.7 on a rerun under rejection pressure. The 180-second per-call abort, 300-second department stall watchdog, and 25-minute run ceiling are all roughly double their measured worst case, with the measurement recorded next to the constant.
The bounce had to defer the design brief
The thirteenth agent ran after the reviewer — which meant on a bounce it built the design brief from the very artifacts the reviewer had just sent back. An accepted bounce now defers it until creative’s revision lands, and the department only reports done once, with the completed output.
Token caps carry headroom, targets govern length
Early caps truncated agents mid-artifact and burned retries. Caps now sit well above the measured output sizes, and the contracts state a 2,000–2,500 word body target — the cap is a safety net, not a length instruction.
How It Is Built
How It Is Built
Application
Next.js 15 App Router, TypeScript, single-page ops console
Models
Claude Sonnet 4.6 on ten agents, GPT-4o as the adversarial skeptic and critic, Claude Opus 4.8 on the press-kit drafter — raw fetch, no SDKs
Contracts
Zod at every department boundary; claims registers inherited verbatim; one corrective retry on invalid agent replies
Streaming
Server-sent events, one EventSource for all runs, heartbeat-driven liveness, reconnect without reload
Canvas
Hand-built SVG in a single requestAnimationFrame loop — eased tweens, particle handoffs, no animation libraries
Fixtures
Incremental capture of every live run; replay at speed multipliers through the same event pipeline the live run uses
Hosting
Railway, Node 20, pinned to one replica for the in-memory run store
What I Built
What I Built
The three-department org: thirteen agent mandates with distinct voices, an execution order derived from declared dependencies, and adversaries that close each department.
The handoff contract: department outputs validated with Zod, claims tagged brief or inferred-with-basis, registers inherited verbatim across boundaries.
The recovery engine: rejects with one engine-bounded retry, the once-per-run bounce with partial rerun and deferred design brief, and a stop control that aborts an in-flight model call cleanly.
The dual-provider router with per-call instrumentation — every call logged with agent, model, elapsed time, and token counts, priced per department.
The console: live SSE canvas in one rAF loop, persistent reject marks, run timers, cost grid, completion summary, and client-side package downloads.
The fixture system: event-by-event capture of live runs, replay at speed with real recorded time and cost, six bundled fixtures including a synthetic bounce rehearsal.
Every part of it built with Claude Code as the development environment, committed after every session.
Honest Limits
What this is not
A system built around marking unverified claims should be equally clear about its own.
The fast demo is a fixture replay of a captured real run, and the console says so on screen — recorded time and cost, labelled as recorded.
Sessions are in-memory on a single replica by design. A process restart loses an in-flight run; completed captures survive as files.
The cross-department bounce is engine-verified and rehearsed by a synthetic fixture, and the reviewer has demanded it against real input — but a full live run has not yet triggered it end to end.
Four of the six bundled fixtures predate the thirteenth agent; they replay correctly with the design-brief slot empty rather than pretending it ran.
Cost figures are per run at current API pricing and will drift as pricing does.
It is a working build and a decision-support demo, not a production platform.
A brief goes in and three departments hand it forward under a contract that carries every claim’s provenance. The adversaries reject what fails, the reviewer bounces what cannot ship, and the number nobody verified never makes it into the layout.