Weekly creates a distinct run with its own cadence, prompt, trace, citations, and draft version.
Operations · Lead Story
Three watch-list signals need operating proof before escalation
A real movement, or a measurement artifact? The investigator desk says probably both. The more interesting story is not that three signals appeared, but that each one stops just short of becoming an operating conclusion.
The five things · what to leave with
- Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze .
- Pine Coast Paid Social softness is plausible, but attribution to the channel is not proven .
- Northeast Fleet may be influencing mix, but current evidence could reflect isolated or temporary work .
- Semantic retry downgraded overclaims and preserved the evidence trail.
- Learning memory now carries forward the proof standards for future Snowflake-backed runs.
Editorial standard
Readable like a reporter. Auditable like a workpaper.
Every material statement is paired with an evidence path. Notes and point comments can add color and hypotheses, but they cannot carry numerical claims unless the source data agrees.
The lead
On the surface, the week looks like a manager's early-warning dashboard: tire pressure is rising, Paid Social looks soft in Pine Coast, and Northeast Fleet revenue is moving faster than bookings. But the agent desk found a more useful pattern underneath. Each signal is visible enough to deserve attention, and each is incomplete enough to punish a premature conclusion.
The result is a briefing that deliberately resists drama. The system is not saying the business is in trouble. It is saying the next operating conversation should be narrower, better evidenced, and easier to prove or disprove. That distinction matters: a weak report would convert movement into certainty; this one turns movement into a testable agenda.
Weekly edition
A slower, more synthetic read for the operating week.
The weekly edition is a separate agent run, not a display mode. It widens the story budget and creates its own draft, trace, citations, comments context, and archive record. Instead of treating every signal as a same-day alert, the agents look for pattern durability, repeat evidence, cross-market differences, and whether prior analyst comments changed the story.
Bring forward notes, comments, prior claims, Ask This Report exchanges, and resolved retries.
Ask what changed, what persisted, and what deserves executive attention across the week.
Weekly draft
Not generated in this session yet
Run Weekly Edition to create a separate draft and archive record.
Proof agenda
What the next query has to settle
Tires
Signal: Appointments and capacity index move together .
Missing: Completed ROs, utilization, cancellations, no-shows, and technician availability.
Decision use: Watch staffing and slot availability, but do not call a capacity squeeze yet.
Paid Social
Signal: Reach holds while appointments soften .
Missing: Spend, impressions, CTR, CPL, lead quality, and appointment conversion by source.
Decision use: Inspect funnel efficiency before changing budget or blaming creative fatigue.
Fleet
Signal: Revenue outruns appointment movement .
Missing: Customer concentration, repeat cadence, labor hours, ticket size, and job category mix.
Decision use: Treat as a mix question until repeat work proves durable demand.
Executive Lead
Evidence points to live business questions, not firm conclusions.
The clean takeaway: management has three signals worth tracking, but the file does not yet clear the bar for escalation . The desk should keep them on watch until conversion, utilization, funnel, and repeat-work data close the gap.
That is a useful management posture. It keeps the agenda sharp without overstating the business case. The next meeting should not ask, “Are tires constrained?” It should ask whether completed work, service capacity, and missed demand agree with the appointment signal.
Tires
Demand signal, not a proven capacity constraint.
Repeated tire-related signals suggest momentum . But appointments and revenue can move for reasons that are not operational capacity: price, mix, scheduling behavior, or work that never converts into completed repair orders.
The operational read is therefore cautious. If completed ROs rise with appointment pressure and utilization tightens, this becomes a staffing and slot-management story. If completed work does not move, the signal may be scheduling friction, pricing, seasonal browsing, or demand that is not converting.
Pine Coast Paid Social
Softness is plausible, but attribution remains unproven.
The file supports a possible Paid Social slowdown . It does not prove channel underperformance without spend, impressions, CTR, CPL, source-level conversion, and appointment trend separation.
The story to test is funnel quality, not vague marketing weakness. If spend and reach are stable while click-through and lead-to-appointment conversion fall, creative fatigue becomes plausible. If spend changed, source mapping drifted, or appointment availability tightened, the channel may be taking blame for an upstream or downstream issue.
Northeast Fleet
Possible mix movement, not yet durable.
Fleet revenue may be influencing mix , but the current support could still be a one-off customer, isolated job, or ticket-size effect. Repeat cadence and customer concentration matter here.
The next query should split the movement by customer, job type, and ticket size. A broad fleet shift would show repeated customers and recurring categories. A fragile signal would concentrate in one account, one work order type, or one unusually large ticket.
Why the retry mattered
The first draft wanted a cleaner story than the evidence allowed.
The semantic retry changed the report's center of gravity. The failed version leaned toward declarative claims: tires constrained, Paid Social underperforming, Fleet mix shifting. The accepted version keeps those ideas as hypotheses and names the tests that would make them publishable.
This is the agentic loop doing visible editorial work. It did not merely polish language; it changed claim strength, preserved citations, and left an audit trail for analysts who want to inspect the rejected path.
Point commentary
Analyst comments stay attached to the claim they are about
Comments are saved as point-level context. In production they would be written to the database with a section ID, paragraph ID, citation IDs, author, timestamp, and revision intent, then retrieved for future reports or draft rewrites.
Run desk
Watch the agentic loop execute
Evidence search
The desk scours, narrows, and preserves the trail
Search queue
Evidence ledger
Hypotheses under test
Retry logic
Where weak drafts are caught before they reach the reader
Rejected overclaim
“Tires are capacity constrained.”
- Missing completed RO proof
- No cancellation/no-show split
- Revenue could be price or mix
Evidence Reviewer sends it back
The retry instruction requires the writer to separate signal detection from causal conclusion, preserve citations, and name the disproof tests.
Rewrite as a watch-list signal. Keep citations. State what would validate or disprove the claim.
Passed editorial gate
“Tire demand signals are present, but completed RO and utilization data are needed before calling a capacity squeeze.”
- Claim strength reduced
- Evidence preserved
- Next checks explicit
Learning mechanism
Run memory becomes habits of success
The system records the causes of retries and successful editorial corrections. Next runs start with those habits loaded into the agent desk.
Lessons carried forward
Next-run behavior changes
- Capacity narratives require completed ROs and utilization before escalation.
- Paid media attribution requires spend, impressions, CTR, CPL, and conversion separation.
- Fleet growth claims require repeat cadence and customer concentration checks.
- Data-quality flags are promoted before any affected metric is used in a headline.
Improvement scorecard
Trace history
Loop history, queries, workpapers, and reasoning summaries
The report can stay clean for readers, while analysts can open the production record behind it: what each agent searched, what it found, what it rejected, and what changed after retry.
Run 01: evidence desk
Retry packet
Rejected draft claim
Tires are capacity constrained; Paid Social is underperforming; Northeast Fleet has shifted mix.
Reviewer reasoning summary
The draft converted weak signals into causal conclusions. The reviewer required completed ROs for tires, source-level funnel metrics for Paid Social, and repeat-customer evidence for Fleet before escalation.
Retry instruction
Rewrite as evidence-bound watch-list signals. Preserve citations. Name what would validate or disprove each claim. Do not use appointment or revenue movement as causal proof.
Learning memory written after run
- Capacity claims must pass completed RO, conversion, cancellation/no-show, and utilization checks.
- Marketing attribution must separate budget, reach, CTR, CPL, and conversion before blaming creative or channel quality.
- Fleet movement must test concentration and repeat cadence before calling a durable mix shift.
- Data-quality blockers are promoted before any affected metric can support the story.
Expandable footnotes
Citations with graphical support
These are the evidence cards the story can lean on. Each figure is drawn with D3 and annotated for the exact point the narrative needs.
[1] Display reporting caveat in Desert ValleyData quality
Display has repeated partial or missing rows, so movement should remain directional and should not become a headline support point.
synthetic_expansion:channel:Desert Valley:Display
0 = ok, 1 = partial, 2 = missing. The annotation highlights why the editor should avoid overclaiming Display revenue.
[4] Tire capacity and appointment pressureService demand
Tire appointments and capacity pressure moved together. The claim still needs completed repair orders before it can become a capacity conclusion.
service_line:Tires:Lakeview:latest_week
The supporting figure shows pressure, not proof. That distinction is the point of the story.
[6] Pine Coast Paid Social funnel softnessMarketing funnel
Impressions can look healthy while appointment yield weakens. The channel story needs source-level funnel metrics before attribution is fair.
channel:Paid Social:Pine Coast:latest_week
The divergence supports investigation, not a final diagnosis.
[8] Northeast Fleet mix movementRevenue mix
Fleet revenue moved faster than bookings, which may imply mix or average-ticket movement rather than broad demand growth.
service_line:Fleet Service:Northeast:latest_week
The desk should inspect RO count, customer concentration, labor hours, and repeat cadence next.
Run archive
Older runs, weekly drafts, and revision history live here
In this demo, run records are saved in browser storage. In production, this page reads from the database: run ID, cadence, prompt, agent trace, citations, draft version, analyst comments, Ask This Report exchanges, retry packets, and publish status.
Analyst notes
Human context that can augment the run
Notes are allowed to shape hypotheses and follow-up checks. The workflow treats them as context, not proof, unless matched to metric evidence.
Deep dive observability
Runtime, handoffs, tools, tokens, retries
Tool calls
load_company_profile, load_channel_rows, load_service_rows, load_analyst_notes, run_deterministic_checks, get_check_results, get_recent_analyst_notes, get_source_summary, compose_briefing
Admin
Schedule routine runs
This hosted demo panel shows the intended operating control: a user can set cadence, instruction, and review mode for scheduled briefings.
About
How this works, and how it becomes enterprise-grade.
What this demo is
This app is a narrative-reporting surface that shows the full agentic loop behind an editorial business report. It demonstrates how a team of specialist agents can search evidence, form hypotheses, run disproof tests, reject weak claims, retry drafts, write fluently, expose citations, and preserve trace history.
Core agent architecture
The intended production workflow uses OpenAI Agents SDK as the orchestration layer. The specialist desk includes Signal Scout, Hypothesis Generator, Investigation Planner, Root Cause Analyst, Evidence Reviewer, Narrative Editor, and Managing Editor. Each handoff preserves evidence, decisions, queries, citations, retry reasons, and learning memory.
Enterprise Snowflake implementation
In an enterprise project, Snowflake becomes the source-of-truth evidence layer. Agents do not freely invent SQL. They call typed tools: source discovery, approved metric queries, validation checks, source-row retrieval, citation creation, trace logging, draft storage, and learning-memory writes. Every material claim links back to Snowflake query specs and source references.
Editable, versioned reports
The production build should store every run, draft, section, citation, query log, trace event, retry event, analyst note, schedule, and memory item. Drafts are editable. Human edits create new versions, preserve the generated draft, and trigger citation-integrity checks when a user introduces new claims or numbers.
Point comments as context
Users should be able to comment directly on a claim, paragraph, chart, or citation. Those comments are saved to the database with their target section and evidence references. Future reports and revision agents can retrieve them as context, so the system learns not only from evals, but also from how analysts actually react to the story.
Clickable SQL evidence
Important claims should carry superscript evidence markers that open the exact SQL/query evidence drawer. The marker should resolve to a citation record containing the query spec, compiled SQL, source rows, freshness status, chart spec, and trace events from the agent that used it.
Retries and learning
Semantic retries catch missing citations, unsupported numbers, hidden data-quality blockers, weak causal claims, and generic writing. After a run, the system records habits of success: proof standards, source reliability, editorial preferences, and failed claim patterns. Future runs retrieve that memory before the agent desk starts.
Operational layer
At enterprise scale, this sits behind authentication, role-based access, scheduled runs, approval workflows, trace redaction, read-only Snowflake roles, query cost limits, eval harnesses, and observability. The reader sees an edition-quality report; analysts can dive into the machinery underneath.
Ask This Report
The conversational layer should stay marginal and grounded. Readers highlight a passage, ask a follow-up, and receive an answer constrained to report context, citations, SQL evidence, analyst notes, comments, and run history. The exchange is stored as another context object that revision agents can use without letting chat replace the crafted narrative.
Page experience
The app is organized as edition pages rather than a single scroll. Transitions should feel like paper being laid down, not a literal page curl: folios update, rules and headlines enter first, evidence unfolds from the margin, and the agent timeline moves like a newsroom production desk.
Embedded comments
Reader commentary can sit beside the claims it changes
These are point-specific comments, not analyst notes. Analyst notes inform the agent run broadly; embedded comments are attached to a claim, paragraph, citation, or chart and can be retrieved by revision agents when that exact point is rewritten.