Context Management & Reliability
CCAR-F · 30 questions
- A long 311 session summarizes aggressively and starts losing dollar amounts and case IDs. What should be preserved separately?
- Verbose work-order tool dumps are bloating the 311 agent context. What should happen before accumulation?
- Aggregated subagent research for a policy brief buries key findings in the middle of a long prompt. What placement helps?
- Downstream context budgets are tight after several research subagents. What should subagents return?
- A multi-issue housing session jumbles IDs, amounts, and statuses across topics. What context design helps?
- A 311 resident explicitly asks to speak with a human. What should the agent do?
- Which triggers should a municipal agent treat as appropriate escalation conditions?
- A city architect wants to route hard cases using sentiment scores and the model's self-reported confidence. What judgment applies?
- A clerk lookup returns multiple residents with similar names. What should the agent do?
- How should a utility replacement agent learn when to resolve a standard meter swap versus escalate exceptions?
- A GIS search subagent fails during a zoning research workflow. What should it return to the coordinator?
- A valid GIS parcel query returns no matching parcels. How should the subagent label the outcome?
- A permit-data subagent hits a transient timeout. What error-propagation pattern is appropriate?
- During multi-source civic synthesis, some topics lack support because sources failed. What should synthesis do?
- Which multi-agent error patterns should a municipal architect reject?
- A sprawling legacy student-information-system exploration keeps losing findings across long sessions. What helps?
- Finding every refund-caller path in a county billing codebase is extremely verbose. What should the main agent do?
- A city codebase agent crashes mid-exploration. How should recovery be designed?
- During a multi-hour municipal codebase tour, context fills with discovery noise. What Claude Code practice helps?
- Before spawning the next wave of exploration subagents, what should the coordinator do with the prior phase?
- A clerk extraction pipeline reports 97% overall accuracy but fails badly on handwritten hardship affidavits. What does this illustrate?
- How should a city monitor error rates even among high-confidence permit extractions?
- Before routing field-level confidence for municipal forms, what should the team do?
- Low-confidence or contradictory fields appear on a digital permit intake. What should happen first?
- When may a municipality safely reduce human review volume on automated extractions?
- Policy research subagents gather citations for a housing brief. What must synthesis preserve?
- Two credible transit reports disagree on ridership. What should synthesis do?
- Open-data findings for a parks briefing omit publication dates. What risk does that create, and what should travel with findings?
- How should a final civic research brief present certainty?
- When synthesizing mixed civic source types, how should financial tables and news articles be rendered?