| AF-AI-001 |
Provider-agnostic AI integration with explicit consent |
supported |
src/ai/ai_service.cpp, src/ai/ai_provider.h |
tests/test_ai_service.cpp |
ADR 0005. No data leaves the machine without user action. |
| AF-AI-002 |
AI profiles and settings |
supported |
src/ai/ai_settings.cpp |
tests/test_ai_settings.cpp |
|
| AF-AI-003 |
OpenAI-compatible provider |
supported |
src/ai/openai_provider.cpp, src/ai/ai_types.cpp |
tests/test_openai_provider.cpp |
Every successful response reports what it consumed (AiUsage: input, cached and cache-written input, output, and the reasoning share of output), or says it reported nothing, so a zero is never read as a free request. An exhausted account is told apart from a rate limit: OpenAI answers both with HTTP 429, and the tool used to report "Rate limit reached" to a user whose only fix was topping up the account (seen during a 300-run evaluation sweep, 2026-09-11); insufficient_quota / credit_balance_exhausted now report "The AI provider account has no credit left". The provider's own error code is kept on every failure. |
| AF-AI-004 |
Local secret storage for API keys |
supported |
src/ai/secret_store.cpp |
tests/test_ai_service.cpp, tests/test_secret_store.cpp |
Windows uses the Credential Manager, macOS uses Keychain Services, and Linux uses libsecret (GNOME Keyring, KWallet and anything else implementing the Secret Service API). Until #53 only Windows had an implementation, so the AI features did not work at all on the other two. test_secret_store.cpp round-trips save/overwrite/delete against the real platform store rather than a mock, which is what makes the Keychain and libsecret calls verified rather than merely written; it skips visibly where no store is reachable. All three backends are exercised in CI: Windows and macOS have a store by default, and the Linux job starts gnome-keyring under dbus-run-session so libsecret is executed rather than only compiled — the step fails if the round trip skips, since gtest exits 0 on a skip and a green step would otherwise prove only that a package installed. Bounded: libsecret is optional at build time, because no keyring API is guaranteed on Linux. A build without it refuses and names the package to install — deliberately, rather than falling back to a file, since a key on disk that the user believes is in a keyring is worse than a refusal. The two ways to be unavailable are reported differently, because they have different fixes: no libsecret in the build says install and rebuild, while a build that has it and finds no running keyring says start one. SecretStoreBackendName() reports which build this is. Release packages: up to 0.3.0-alpha.2 the release workflow's Linux job did not install libsecret-1-dev (CI's did), so the published Linux archive was a build without libsecret that refused to store a key — checked on 0.3.0-alpha.2, whose binary carries the "this build has no keyring support" message and no libsecret dependency. The job installs it from 0.3.0-alpha.3. |
| AF-AI-005 |
Background AI task execution |
supported |
src/ai/ai_task_runner.cpp |
tests/test_ai_task_runner.cpp |
|
| AF-AI-006 |
SCCG-guided AI element review |
supported |
src/review/sccg/sccg_review.cpp, src/review/sccg/sccg_review_preparation.cpp, src/review/sccg/sccg_review_passes.cpp, src/review/sccg/sccg_profile_selector.cpp, src/app/actions/ai_review_actions.cpp, src/app/actions/proposal_actions.cpp, src/app/controllers/ai_review_controller.cpp, src/app/app_runtime_ai.cpp, src/ui/gsn/gsn_badges.cpp |
tests/test_ai_claim_review.cpp, tests/test_sccg_review_preparation.cpp, tests/test_sccg_review_passes.cpp, tests/test_ai_review_controller.cpp, tests/test_ai_review_actions.cpp, tests/test_ui_state.cpp |
One AI Review action selects the unique SCCG profile for the selected element's element role — claim, strategy, evidence, assumption, justification, context, or challenge — which is the notation-neutral key SCCG publishes for a tool whose model is neither GSN nor CAE. The same key names the selected-element data package the element is sent in, looked up from the profile rather than hard-coded, so the profile chosen and the package sent cannot disagree. It reviews the complete integrated working draft, not only accepted SACM. Suggested corrections become independently reviewable SCCG draft groups linked to their finding and guideline. A finding may now propose a structural repair where that is what SCCG prescribes -- creating a Strategy for AR.2, a Solution for EV.1, a Context or Assumption for AR.3/AR.6/AR.7 -- staged as one group so a reviewer accepts the repair or none of it, rather than only rewording the reviewed element. The request now carries the case's own glossary and any prior findings on the elements under review (including findings on the parent and children it was shown, not only the selected element), and every absent data package says why it is absent -- the tool has no source, the case holds none, or it was deliberately withheld -- so a judgement bounded by what was not shown cannot read as a judgement that the data does not exist. A proposal may touch only the elements the review was shown (the profile's own data packages) and elements it creates itself; anything reaching further is refused and reported. A result is refused only if the elements it reviewed changed while the request ran -- an edit elsewhere in the draft, by the user or another contributor, no longer discards it. A successful no-findings result persists as a green check badge after the spinner stops. Selection fails closed if the catalog has no match or more than one match. An absent package now carries SCCG's own instruction for reviewing without it where the profile publishes one: evidence_review says to report the citation and control findings, state that sufficiency was not assessed, and not to report the absent basis as a finding against the argument, naming the eight guidelines (EV.5, EV.6, SU.3, SU.6–8, LF.5, LF.7) that cannot be assessed without it. That turns the degradation every evidence review already ran under from an improvised one into a declared one. The three absence states are SCCG's published availability_states rather than this tool's own names. Correction: until this row cited sccg_review_preparation.cpp, only three packages (PROJECT_GLOSSARY, CHANGE_HISTORY, USER_REVIEW_INTENT) ever reported empty; every other package the tool builds — PARENT, CHILDREN, DIRECT_CONTEXT, INHERITED_CONTEXT, STRATEGY, EVIDENCE_PATH, the selected-element packages — fell through a catch-all that declared them not_implemented. A root goal's absent parent and a claim's absent context were therefore reported as data the tool cannot produce, when the fact was that the argument does not contain them. Those two read in opposite directions: the first excuses the argument, the second is a finding against it. STANDARD_LINKS remains genuinely not_implemented — it has a required field (linked_requirements) this tool cannot fill. Under SCCG 0.8.0 an unavailable package silences nothing. EVIDENCE_BASIS has no source here and is reported not_implemented; 0.7.0's evidence_review when_absent statement turned that into eight unassessable guidelines (EV.5 measured 0 of 3 against evidence stating no sufficiency basis), and this row once recorded a workaround that sent the package available with every field empty. 0.8.0 fixed the cause (safety-case-core-guidelines#13) and defined a package with no required fields and nothing in it as empty, which made the workaround non-conforming, so it is gone. SCCG's when_unavailable rule and its meaning for each availability state now reach the model verbatim, replacing this tool's paraphrase. Under SCCG 0.9.0 the availability rule is uniform: a package is available when its required fields are present and at least one of its fields is populated, whatever it requires, so a claim's CHILDREN sent as {"child_elements": []} is reported empty rather than available-with-nothing (safety-case-core-guidelines#17). Correction: only the second half of that rule was checked, so a package with a required field left out but another filled counted as available; it is now reported not_implemented, naming the missing fields -- not empty, since the case may hold what was left out. 0.9.0's when_unavailable also lets the review rely on an empty package as a fact about the case -- a claim with no children has no path to evidence -- while not_implemented and withheld still tell it nothing. Every package uses SCCG's published field names, since 0.8.0 judges availability on them: USER_REVIEW_INTENT was sent as intent and CHANGE_HISTORY as review_items, neither of which the catalogue names, so both read as empty by its definition; history is now filed as prior_findings (AI) and review_comments (people), and a test checks every package the collector builds against the registry. A recorded finding citing a retired guideline keeps its id, with SCCG's redirect beside it. claim_review is sent as its four published review passes, concurrently, each carrying only its own guidelines and framed by SCCG's published review_pass_instruction sent verbatim with the pass's question substituted (0.9.0; this tool's own framing sentence is kept only for a catalogue that predates it; 0.10.0 added the rule that a defect a distinguish_from note assigns to a guideline outside the pass is not a finding in that pass, which reached every pass request with no code change), merged by MergeReviewPasses: a finding cited outside its pass is discarded and reported, and a failed pass makes the review incomplete — its findings recorded, an "AI review incomplete" item naming the pass, the outcome failed so no badge is earned. Measured on the probe corpus, splitting took 18 of 31 claim guidelines cited as intended to 26, none worse. It costs input: 195,338 bytes against 133,176 for one request on the Paper A top goal (+47%), because packages and instructions repeat per pass. Each guideline also carries SCCG's distinguish_from notes and its own tool.repair, and the response contract translates SCCG's repair vocabulary into operations once, replacing a hand-kept per-guideline list that named three guidelines SCCG has since retired. Steps 1–5 of the workflow now live in review::PrepareSccgReview, so the request the application sends and the request the evaluation harness (AF-AI-026) sends are assembled by one implementation. The response contract no longer asks for a severity. SCCG defines no severity concept for a guideline, so the field was this tool's own invention, and the prompt told the model it "should normally be warning" — which is what it got: 149 findings out of 149 marked warning across a 57-run measured sweep, a field carrying nothing a reviewer could rank or filter by. Severity is now assigned by the tool (every SCCG finding is an observation for a human to judge). Ranking uses two signals the model or the tool can actually establish: confidence, which the contract now defines against the supplied data rather than leaving undefined (and which does vary — 123 high to 26 medium over the same sweep), and pre-check corroboration via CorroboratingPrecheckIds, which names the deterministic checks that independently reached the same guideline. |
| AF-AI-007 |
MCP server — read a safety case from an external AI client |
supported |
src/mcp/server.cpp, src/mcp/tools.cpp, src/mcp/session.cpp, src/agent/read_operations.cpp, src/app/agent_request_handler.cpp, src/app/mcp_client_config.cpp |
tests/test_mcp_server.cpp, tests/test_mcp_modes.cpp, tests/test_mcp_reconnect.cpp, tests/test_agent_request_handler.cpp, tests/test_mcp_client_config.cpp |
Design: docs/features/mcp-server.md. assurance-forge-mcp speaks JSON-RPC 2.0 over stdio and reads a case (overview, search, element detail, GSN tree, argument files). It runs in one of two modes, re-evaluated on every call. Connected: Assurance Forge answers from the complete integrated working draft the user is looking at, including MCP, SCCG and human groups. Every case-content result names view, argument_file, workspace_id and working_revision, so an agent never mistakes unaccepted text for accepted SACM. Offline: no application is reachable, so the adapter reads accepted SACM from its own copy and reports view: accepted; it never opens .af/drafts and remains read-only. Both modes execute the same src/agent/ read implementations. The session heals without restarting the client: an offline session promotes itself when the application appears, a lost connection reconnects on the next call (the application restarting or switching projects is an ordinary event, not a dead session), an interrupted read is retried once after reconnecting, and an interrupted mutation is reported but never replayed — the application may have applied it before the connection broke. get_connection_status reports the current mode, application version and consent state without returning case content, so it works before consent is granted; initialize states the mode in its instructions. The session id survives reconnects, so draft-group ownership persists across an application restart. Consent is two gates: the ADR 0007 master flag (fails closed, re-read on every call, verified through the real process by cmake/run_mcp_smoke_test.cmake) and the ADR 0014 per-session access grant -- the user approves each session's access to the open project in the application, and ungranted operations are refused with project_access_pending. Independent of AF-AI-001..006: the layer gate forbids mcp/ from including ai/. |
| AF-AI-008 |
Local bridge between the MCP adapter and the running application |
supported |
src/bridge/protocol.cpp, src/bridge/transport.cpp, src/bridge/instance_registry.cpp, src/app/controllers/agent_bridge_controller.cpp, src/app/agent_request_handler.cpp |
tests/test_bridge_protocol.cpp, tests/test_bridge_transport.cpp, tests/test_instance_registry.cpp, tests/test_agent_bridge_controller.cpp, tests/test_agent_request_handler.cpp |
A versioned, user-private, message-framed connection: Windows named pipe, POSIX AF_UNIX socket, no TCP. The application publishes one instance record per running instance in the user's runtime directory (ADR 0014) — never in the project, which keeps af.proj under a single writer — carrying a fingerprint of the open project, not its path. The listener outlives project open/close/switch; a session bound to a project that is no longer active is refused with project_not_active rather than shown the new project, and a second instance opening an already-open project gets an advisory warning. Records are pruned by pid liveness. Two gates: the operating system's access control on the pipe (an explicit user-only DACL on Windows, 0600 on POSIX), and a 256-bit token from the record. Requests execute on the frame thread, so an operation sees exactly the model the user sees and needs no lock. Every frame carries a protocol version, and a mismatch between an old adapter and a newer application is reported as one message naming both versions. Both platform branches are compiled in CI; the Windows branch was additionally built with MinGW g++. |
| AF-AI-009 |
Propose changes to a safety case from an external AI client |
supported |
src/core/drafts/draft_document_store.cpp, src/core/drafts/draft_operation_apply.cpp, src/core/drafts/draft_workspace.cpp, src/core/drafts/draft_workspace_store.cpp, src/agent/draft_operations.cpp, src/app/agent_request_handler.cpp, src/app/app_runtime_project.cpp, src/mcp/tools.cpp |
tests/test_agent_draft_document.cpp, tests/test_agent_request_handler.cpp, tests/test_draft_workspace.cpp, tests/test_draft_operation_apply.cpp, tests/test_mcp_server.cpp, tests/test_mcp_modes.cpp, tests/test_patch_operation_parsing.cpp |
An agent contributes attributed change groups to the one persisted working draft for the argument currently open. It can begin, stage, replace, inspect, submit, remove and close its own groups, describe the combined draft, and poll revisioned events. Every modifying call carries expected_working_revision and expected_context_generation; a human, SCCG or other MCP mutation fails an old call with current_working_revision, and a project switch, fresh grant or revocation fails it as stale context, before anything is stored. Stable generated ids let a later call develop an element an earlier call created, and connected reads see that element because they use the working model. An MCP session may inspect every contribution but cannot mutate another contributor's group. Staging writes only .af/drafts recovery state — no accepted SACM, command or audit transaction — and survives restart with provenance and identities intact. Promotion stays the ordinary audited ApplyProposalCommand exposed only to the human UI. A registry-wide test proves there is no apply, accept or promote tool. Change-set tool names temporarily alias the group operations and no longer create private ChangeSetStore state. Staged operations are now applied to the draft SACM document (ADR 0016), through the same library seams the application uses on the accepted document, so an operation the model cannot hold is refused in the call that made it and named by its position in the batch — rather than staged, drawn on the canvas as pending, and refused at accept. Batches are atomic: a refused batch leaves the draft exactly as it was. created_element_ids now carries ids the document allocated, addressable immediately, because no later materialization can reallocate them. replace_change_group, remove_change_group and unstage_operations are refused against a document-backed draft and say what to do instead: they withdraw operations from a log, and reporting them as done while the draft still held the change would be the silent-drop defect in a different call. expected_working_revision now tracks the draft document, the counter that moves whenever anyone edits the argument — including the user, whose edits go into the same draft — so a client call computed before such an edit is refused. Support attaches by the relationship the child's GSN role requires, through the shared apply_attach_child seam: a Strategy is the reasoning of an inference rather than one of its ends (its inference is deferred until a sub-goal gives it a source, SACM clause 11.13), and a Solution attaches by AssertedEvidence — the only thing that distinguishes it from a Context, which is the same SACM type. Getting either wrong produced an argument that drew correctly and failed the GSN well-formedness check, or a Solution rendered with Context notation. An operation carrying a key the parser does not read is refused, naming the key and where its value belongs. It used to parse cleanly and report success: a CreateTerm sent with "definition" rather than "new_value" produced a term with an empty definition, so a glossary an agent had been asked to write arrived with every word undefined and nothing anywhere saying a definition had been dropped. A key the parser does read but whose value is the wrong type -- "new_value": 123 -- is refused for the same reason: it reached the applier as an empty string, so the operation applied nothing and still reported success. |
| AF-AI-010 |
SCCG guidance and checks for an external AI client |
supported |
src/mcp/guidance.cpp, src/core/sccg/staged_checks.cpp, src/core/guideline_catalog.cpp, src/agent/draft_operations.cpp |
tests/test_sccg_staged_checks.cpp, tests/test_mcp_server.cpp, tests/test_agent_request_handler.cpp, tests/test_draft_workspace.cpp |
Three mechanisms. Resources: sccg://guidelines publishes the catalog over MCP, opening with the title, purpose, version and licence from the catalog's own document block — the hardcoded heading fallback is gone, because since SCCG 0.7.0 the dist files the runtime loads carry that metadata too and no longer only the YAML fallback did — and resources/templates/list publishes sccg://guideline/{id} for one guideline at a time. The MCP executable carries its own data/sccg/dist copy, and the stdio smoke test drives resources/read and prompts/get through the real process in both consent modes — SCCG is the public corpus, not case content, so it is deliberately readable before consent. Prompts: draft_argument_from_standard, add_argumentation, restructure_case and translate_case each carry the guidance for that job, quoted from the catalog so prompt and guideline cannot drift (translate_case quotes CL.5, the qualifier rule its prose states informally); the workflow text tells agents to use revision-checked integrated draft groups, and to write every new element in each language the case is maintained in (AF-AI-019). Checks: every staging result returns SCCG and structural findings against the complete materialized draft, and the same findings are available through describe_working_draft. The mechanical set is the deliberately narrow, individually tested subset AF-AI-024 records — thirteen checks, each bound to the catalog's own check id and enumerated by core::sccg::ImplementedCheckIds(), which is the same list every result now reports. Most of SCCG is prose only a reader can judge, so this is not "SCCG compliance": findings are advisory and never block promotion. Every result now says which checks it could decide and states plainly that no findings is not conformance, and every finding carries the guideline's own wording plus its sccg://guideline/<id> resource, so an agent holds the rule rather than the tool's paraphrase of it. The authoring prompts no longer carry a hand-picked guideline list: each names the SCCG element roles its output produces, resolves those to review profiles through the catalog's published authoring_guidance.element_rules, and quotes their guidelines, so the criteria an agent writes to cannot drift from the criteria applied to it — the role-to-profile mapping was the last part of it held in code. All three mechanisms here land only when the client or user reaches for them; AF-AI-023 adds the channels that land unprompted, and docs/features/mcp-authoring-quality-plan.md is the plan for widening the whole surface. |
| AF-AI-011 |
Suggest where new argument belongs |
supported |
src/agent/placement.cpp |
tests/test_mcp_modes.cpp |
suggest_placement(topic) returns ranked goals and strategies with the path from the top goal, the sub-claims already there and the context in scope. Substring search says where a word appears, which is a different question from where an argument belongs; without this an agent guesses and attaches at the root. Ranked by term overlap and structural fit, not by understanding — the reply says so, and says to report that nothing fits rather than pick the best of a bad set. Only goals and strategies are offered as anchors. |
| AF-AI-012 |
Atomic re-parent for restructuring (MoveUnder) |
planned |
|
|
Restructuring is expressible today as remove-then-add supported-by pairs, so the capability exists, but a large restructure reads as a pile of unrelated operations and the intent is lost in the diff a reviewer has to approve. A MoveUnder patch operation would carry the intent atomically. Adding a PatchOperationType is backward compatible for reading existing audit logs. Not built. |
| AF-AI-013 |
Standards-clause traceability on claims |
candidate |
|
|
An agent applying a standard records which clause a claim answers in the claim description, because the element model has no citation field. That is a trace a human can read but not one a tool can follow, filter or check for coverage. A first-class citation would be a SACM modelling decision, not an MCP one. |
| AF-AI-014 |
One integrated working draft per argument file |
prototype |
src/core/drafts/draft_workspace.cpp, src/core/drafts/draft_workspace_store.cpp, src/core/drafts/draft_document_store.cpp, src/core/audit/audit_accept.cpp, src/agent/draft_operations.cpp, src/app/agent_request_handler.cpp, src/app/actions/ai_review_actions.cpp, src/app/actions/proposal_actions.cpp, src/app/app_runtime.cpp, src/app/app_runtime_project.cpp |
tests/test_draft_workspace.cpp, tests/test_agent_request_handler.cpp, tests/test_ai_review_actions.cpp, tests/test_audit_accept.cpp |
One persisted workspace per argument materializes ordered changes from MCP, SCCG AI review, and human draft editing over accepted SACM. The draft is projected through the same render passes as the accepted argument, so a term defined in a draft renders with its definition rather than as an undefined context node -- the draft view was a bare projection, so a term an agent had just defined showed on the canvas as though it had none until a restart re-ran the passes through load_file. Connected MCP reads and SCCG review both consume that combined model and write attributed groups back into it; revision checks prevent stale MCP calls from overwriting intervening contributors. Accepted SACM remains unchanged until human promotion. Kept at prototype while the merged workspace UX and end-to-end source combinations receive stability testing. The draft is now a SACM document rather than a list of operations (ADR 0016): contributors edit it through the library seams, the working view is its projection, what it changes is a comparison against the accepted document, and accept is one atomic write of that document over the argument. A draft is created by the first unaccepted change rather than when an argument is opened — one cloned per argument opened stopped descending from the argument it is compared against the moment the user edited the accepted document, and reported their own new elements as removals. While a draft exists the user's own edits go into it, for the same reason — argument edits and glossary edits alike, the latter expressed as the terminology operations an MCP client sends. The change-group ledger survives beside it recording who contributed what, and an argument the SACM library cannot load still falls back to the operation-staging path. The accept is recorded in the audit log as one AcceptWorkingDraft transaction whose WorkingDraftAccepted event carries the accepted document in full and the draft's provenance (groups, sources, guidelines, rationales). The event is replayable, so the history slider reconstructs the states on either side of an accept, and the accept is promoted to the trusted replay root with a snapshot at its own sequence, so the next project open verifies instead of reporting the accepted .sacm as a divergence (#409). Reported from a demo project built over MCP: everything the agent contributed was in the file and absent from the log, and the only remedy offered archived the history. The snapshot makes the accept an undo boundary — the draft it consumed is gone, so Ctrl+Z stops there and restore-from-history is the way back. Correction: while the argument drafted as a document, SCCG review did neither of the things this row claimed. It read the change-group materialization, which holds only what MCP recorded, so it judged accepted wording the user had already replaced in the draft; and its suggestions were staged into change groups only, so Accept wrote the document without them and then cleared their groups. Both now use the draft document. Provenance travels with the element as assuranceForge.draft.<contribution>.<field> tags written in the same all-or-nothing batch as the change, by MCP, SCCG review and the user alike, and stripped at accept (#409); a removed element carries none. See ADR 0009, ADR 0010, ADR 0016 and docs/architecture/integrated-draft-workspace-plan.md. |
| AF-AI-015 |
Dependency-aware selective promotion of draft changes |
prototype |
src/core/drafts/draft_dependency_graph.cpp, src/core/drafts/draft_promotion_service.cpp, src/app/app_runtime_project.cpp, src/ui/panels/element_panel.cpp |
tests/test_draft_workspace.cpp |
Accept a coherent change group without accepting the rest of the draft, with the dependency closure computed and shown first: a wording edit promotes alone, a new argument branch cannot promote without the strategy and relationships that make it meaningful. Remaining groups are rebased onto the prospective baseline and validated before anything is written, and promotion is refused before mutation if they cannot be. Promotion itself stays the unchanged audited ApplyProposalCommand, and undoing one restores the accepted baseline and the pre-promotion draft together (AF-AI-017). The deferred library re-derive bumps case_revision when it swaps the models, so the per-package canvas tab rebuilds: the dispatch bumps that counter a frame earlier while the old projection is still in place, and without a second bump the tab matched its own stamp and kept drawing text the accepted change had already replaced while the inspector showed the new text beside it. A promotion that fails after its marker is written clears the marker rather than leaving the workspace in Promoting, where it refuses editing, accepting and discarding until the application restarts — a failed accept must not take away every way to respond to it. The surviving groups are re-anchored to the argument the library produced, not to the one the patch predicted: the two agree for argument edits but not for terminology, where the seams stamp a gid a flat patch cannot know, and anchoring to the prediction declared every surviving group stale against the argument the user had just accepted into. Selection is per group from the Draft Changes panel (AF-AI-018) or per element from the Inspector's contribution list, with accept-all from the banner. Rejecting a group that others are built on now names them and offers the choice: reject them too, or keep them marked NeedsAttention — excluded from materialization, because left in they would fail to apply and block the whole draft, but retained in the workspace and recoverable by retargeting their operations. The audit transaction carries the promoted group ids, the contributing source labels, the guidelines served, the review items answered and each author's rationale, because accepting a draft consumes it and the log is then the only record of where the change came from. Kept at prototype until the release-gate scenario in #273 passes in the running application. Does not apply to a document-backed draft (ADR 0016), where accept is all-or-nothing: there is no selection of operations left to compute a closure over. Selective review inverts instead — the reviewer removes what they do not want from the draft and accepts what remains — and the gesture for doing so in the UI is not built yet, so a reviewer who wants only part of a document-backed draft currently has to edit it down by hand (#409). |
| AF-AI-016 |
Draft recovery across restart |
prototype |
src/core/drafts/draft_persistence.cpp, src/core/drafts/draft_workspace_store.cpp, src/core/project_service.cpp |
tests/test_draft_workspace.cpp |
Unaccepted work survives closing the application, stored under .af/drafts/ and reconstructed from the accepted baseline plus its change groups — never as SACM, which has no assertion state meaning "AI-proposed and not accepted". A stored draft whose base hash no longer matches the argument enters NeedsRebase and replays nothing. Promotion writes a pending marker before it touches accepted SACM, so a crash mid-promotion finalizes or cancels from the accepted-model hash rather than applying twice. .af/ is generated with a .gitignore so unaccepted AI-authored argument text does not reach a colleague through version control. Kept at prototype with AF-AI-014: recovery is only as trustworthy as the workspace lifecycle around it, which is still under stability testing. |
| AF-AI-017 |
Undoing an acceptance restores the draft it consumed |
prototype |
src/core/drafts/draft_persistence.cpp, src/core/drafts/draft_workspace_store.cpp, src/app/app_runtime_project.cpp, src/app/app_runtime_undo.cpp |
tests/test_draft_workspace.cpp |
Promotion is one boundary on the accepted undo stack, but a draft is deliberately not a command and has no entry on that stack. Undoing a promotion therefore took the change out of the accepted argument while the draft had already given it up — and where the promotion consumed the last group, the workspace was deleted with it, so the work was in neither place. Promotion now records the pre-promotion workspace under .af/draft-promotions/<transaction>.json, outside the per-argument draft directory the acceptance deletes, and undo puts the groups back with their provenance and generated identities, rebased onto the model the undo restored. Groups staged after the promotion are kept rather than replaced. A snapshot that cannot be read, or that belongs to another argument, refuses the undo instead of destroying the only copy of the work. Undo is now two stacks. While a draft holds unaccepted edits, Ctrl+Z reverses the last one and the accepted stack is untouched; it falls through to the accepted history once the draft has nothing left, which is what keeps a promotion undoable while later groups are still staged. The workspace revision still moves forward across a draft undo, so a token minted before it is not silently revalidated by content that happens to match, and next_sequence is monotonic so an undone group id never names a second group in one event log. A promotion clears the draft undo history: every entry below it describes groups the accepted argument now contains. The history is session state, not recovery state — the draft's content is persisted, so a restart recovers the work with nothing left to undo. Promotion snapshots are pruned against the audit undo boundary, the only rule that cannot delete one an undo could still reach. |
| AF-AI-018 |
Draft Changes panel — every unaccepted change, whatever wrote it |
prototype |
src/ui/panels/draft_changes_panel.cpp, src/app/areas/draft_changes_area.cpp, src/app/areas/feedback_dock_area.cpp, src/app/areas/canvas_history_overlay.cpp |
tests/test_draft_changes_panel.cpp |
One row per change group in the working draft, replacing the split proposal / change-set views that showed one source at a time and could not show a combination at all. Each row names who wrote it and in which session, its rationale, what it adds, changes and removes — with relationships counted separately from elements, because a changed support relationship can alter the meaning of an argument more than a reworded claim can — the guidelines it serves, the review items it answers, what it depends on, findings against what it would produce, and whether it can be accepted right now. Promotability is answered by planning the promotion against the materialized working model rather than guessed at, and a row that cannot be accepted says so on the row rather than in the status bar. Accepting names what else it would accept before the button is pressed. Selecting a row takes the user to its first changed element in whichever view can show it: the GSN canvas for argument, the terminology view for a term that is already in the accepted glossary, and nowhere at all for a term this draft created — that view reads accepted terminology, so going there would report the term missing, and the row's own glossary lines are where it is readable. A glossary group additionally lists each staged term with its definition, categories and source in full on the row — a term is deliberately not a GSN node (AF-AI-022), so there is no canvas rendering beside the row to read it from. A whole draft held back — stale, promoting, or unmaterializable — is explained once above the list, because in that state no single group is the explanation. An accept that refused is explained the same way, on the banner beside the button that appeared to do nothing and above the list, until the draft changes: previously the only report was the status bar, which is one line and truncated the sentence before the reason, so a user who pressed Accept all was left with a banner still counting unaccepted changes and nothing on screen saying why. Prototype: not yet seen in the running application against a real multi-source draft, which is the #273 release gate. |
| AF-AI-019 |
Bilingual argument from an external AI client |
prototype |
src/core/reviews/review_proposal.cpp, src/core/reviews/review_proposal_patch_service.cpp, src/core/reviews/review_proposal_plan.cpp, src/agent/change_operations.cpp, src/agent/read_operations.cpp, src/core/drafts/draft_promotion_service.cpp, src/mcp/tools.cpp, src/mcp/guidance.cpp, src/app/app_runtime_project.cpp |
tests/test_review_proposal_patch_service.cpp, tests/test_draft_workspace.cpp, tests/test_agent_request_handler.cpp, tests/test_change_set_acceptance.cpp |
An agent states each element in every language the case is maintained in, within the one operation that creates it: text/new_value carries the primary language and translations carries the rest, so a reviewer accepts a bilingual claim or none of it and no group promotes half-translated. An UpdateElementText carrying only translations revises those languages and leaves the primary text alone, which is what makes translating an existing argument safe — translating a safety case must not edit it. A Create* with translations and no primary text is refused: that element would render empty for every reader who has not switched languages. Reads report translated_languages per element and the text under translations, and find_elements matches either language. The same vocabulary unblocked human secondary-language edits while a draft is active, which were previously refused outright. Translations from a non-human source arrive flagged TranslationReviewNeeded on promotion (AF-ENG-012): accepting the argument is not the same as establishing that the Japanese says what the English says. The element semantic hash now covers secondary-language text, so a proposal written against an untranslated element no longer looks current after someone translates it. A bilingual group could be staged, shown and planned but not accepted: the proposal planner declined a non-primary-language name as unrepresentable, three hours before the adapter gained the reserved sacm.import.name write that carries exactly that (AF-ENG-012, SACM23-LIB-002), and the rationale was never revisited. With the compatibility path since removed, Accept All refused an entire 64-operation draft over one translated goal name — a person could type that same name into the inspector and it saved. Acceptance of a bilingual group, including a created element's translations, is now asserted end to end through the promotion plan, the preflight and the saved file. Prototype: the write path still rides the draft workspace under stability testing (AF-AI-014). |
| AF-AI-020 |
Assurance claim points readable from an external AI client |
supported |
src/agent/read_operations.cpp, src/mcp/tools.cpp, src/app/agent_request_handler.cpp, src/mcp/session.cpp |
tests/test_agent_acp_reads.cpp, tests/test_assurance_claim_point.cpp |
list_assurance_claim_points serializes the projected ACP records: what each one annotates — an element such as a Solution, or a SupportedBy/InContextOf relationship — and how it is resolved: inline text, or a confidence argument whose claim, package and top-goal ids are followable with get_element; an ACP with no resolution reports instantiated: false, the state the ACP panel warns about. get_element carries the ACPs on the element and on the relationships touching it, relationship-borne entries naming the relationship they ride. Available connected (working-draft view) and offline. Read-only, deliberately: the patch vocabulary has no ACP operation, and extending it is an ADR 0009 vocabulary decision with its own GSN review — authoring stays in the application (AF-ACP-001..008). ACPs seen through the working draft come from the accepted baseline, since no draft operation can create one. |
| AF-AI-021 |
Projectless MCP sessions with runtime project binding |
prototype |
src/mcp/session.cpp, src/mcp/main.cpp, src/app/controllers/agent_bridge_controller.cpp, src/app/areas/modal_host.cpp |
tests/test_mcp_dynamic.cpp, tests/test_agent_bridge_controller.cpp |
Launched with no project argument, the adapter initializes without a running application, discovers the single running instance at call time, and connects unbound (ADR 0014). Access is granted per session: the first project operation (or request_project_access) raises an in-app request — client label, project, Allow while open / Deny — and is refused with project_access_pending until the user answers; this applies to --project sessions too. Grants are keyed by session id, survive a reconnect to the same instance, and end on deny, revoke, project close/switch, MCP disable, or restart; a re-granted session still owns its draft groups. Multiple running instances are never auto-selected. --offline-project <path> is the deliberate read-only escape hatch that never connects. Prototype: the single-instance path is complete and tested end to end through the real adapter process; the setup UX is not. With more than one application running the session refuses and names the fix rather than choosing, and the client configuration Preferences copies always pins --project <the open project>, so the projectless connect-once setup this row is about still has to be written by hand -- both are expected to change how a user first connects. (Context envelopes and generation checks shipped with AF-AI-009; rebinding is the fresh grant raised after a project switch.) |
| AF-AI-022 |
Terminology read and management from an external AI client |
prototype |
src/agent/read_operations.cpp, src/agent/change_operations.cpp, src/core/reviews/review_proposal_patch_service.cpp, src/core/reviews/review_proposal_plan.cpp, src/sacm_adapter/document_edit.cpp, src/core/commands/proposal_commands.cpp, src/mcp/tools.cpp, src/core/drafts/draft_operation_apply.cpp, src/app/areas/workbench_area.cpp, src/app/actions/terminology_actions.cpp, src/ui/panels/terminology_package_panel.cpp |
tests/test_mcp_server.cpp, tests/test_agent_request_handler.cpp, tests/test_review_proposal_patch_service.cpp, tests/test_change_set_acceptance.cpp, tests/test_draft_operation_apply.cpp, tests/test_term_definition_survives_save.cpp, tests/test_working_glossary.cpp |
list_terms returns every term's value, name, definition, categories, external reference and origin in one call, together with the categories the case defines, connected and offline; terms and categories also appear in get_case_overview counts, find_elements and get_element. CreateTerm, UpdateTerm, RemoveTerm, CreateCategory and UpdateCategory stage through the same revision-checked change groups as argument edits and promote through the same audited ApplyProposalCommand; accepting the first term or category of a case with no glossary creates the containing terminologyPackage rather than refusing. UpdateTerm fields are value, definition, name, category (space-separated ids), external_reference and origin — the last three are what answer the terminology check's "no category" and "no external reference/source" findings, which an agent could previously read but not fix. Each field is written by its own seam so classifying a term cannot rewrite its definition or drop the translations of it; an unresolvable category or origin is refused at staging, not at acceptance. Definitions may carry translations; a term's value is a single string (SACM 10.11) and staging refuses a translated one with an explanation. Addresses SCCG CL.5 by defining a bounding term once instead of repeating it as free text. Bounded: removing a category, and associating a term with an element as a visible context, stay in-application (a category delete has cascade semantics needing a confirmation this surface cannot raise); created terms and categories land in the case's first terminology package; removing a term an argument package references is refused at acceptance by the library's cross-package delete guard. Prototype with AF-AI-014: the write path rides the draft workspace, which is still under stability testing; the read path is exercised through the real MCP server process. A CreateTerm with no definition is refused, on the staging path and the review path alike, and list_terms reports how many existing terms have none. Reported twice from real sessions: an agent staged a whole glossary, set the category and external reference the guidance names as checked, and left every definition empty — because nothing asked for one. The terms reached the accepted argument as words with nothing beside them, which reads as a defined glossary and is not. The definition itself was never the broken part: it stores and serializes through both the create and the UpdateTerm path, and there are now tests pinning that. A term is matched against element text by its value, so the schema now says the value must be the string exactly as it is written in the argument — EPB, not EPB (Electronic Parking Brake) — with the expansion in the definition, and list_terms marks every term whose value appears nowhere in the case. All four abbreviations in a reported session were staged with the expansion inside the value, so they bound nothing and every occurrence read as undefined. The working glossary is what the application shows (ADR 0016): the terminology tab, the canvas's term detection and the terminology checks read the draft document's package while it differs from the accepted argument, every row the draft added or changed is badged draft with the fields that differ, a notice above the table counts the unaccepted glossary changes, and a glossary the draft itself created is shown although the accepted package has none. Reported from a demo: a definition an MCP client revised was invisible until a restart, because every terminology surface read the accepted package. The user's glossary edits go into the draft too: while a draft document exists, a term or category added, changed or deleted in the terminology tab (and a term defined from the canvas) is applied to the draft document as the same CreateTerm / UpdateTerm / RemoveTerm / CreateCategory / UpdateCategory operations a client sends, one per changed field, through the same seams — so a human edit is accepted or refused exactly as a client's is, the accepted glossary and its audit log stay untouched until the accept, and the tab says where the edit goes before the first click. Previously those edits were refused outright, because they wrote to the accepted document the draft no longer descended from. Bounded: what the vocabulary cannot express — the package's own name and description, deleting a package or a category, linking a term to an element as context — is still refused with a reason while a draft is open, and a new term or category lands in the case's first glossary. |
| AF-AI-023 |
Authoring doctrine delivered in every MCP session |
supported |
src/mcp/guidance.cpp, src/mcp/server.cpp, src/mcp/tools.cpp |
tests/test_mcp_server.cpp |
Phase 1 of docs/features/mcp-authoring-quality-plan.md. The SCCG prompts and resources (AF-AI-010) land only when the user asks for them; this row is the answer for the user who types "add an argument that braking is safe" and says nothing about SCCG. The authoring doctrine — one rule per line, each naming the guideline it condenses — travels in the three channels that reach the model unprompted. It is rendered from the catalog, not written here: SCCG 0.7.0 publishes authoring_guidance.core_rules, the eighteen-guideline subset a tool should deliver while an author is writing rather than when a review is run, each with a one-line short_rule and a recorded reason for inclusion. The fifteen lines this used to maintain by hand are gone, and with them the "every SCCG family is represented" property that comment asserted on its own authority — a test now checks it against the published subset. The rendering follows the file's own usage: from short_rule, citing the id, without paraphrase, and carrying SCCG's caveat that the subset is a delivery subset and not a reduced standard. The three channels are: initialize.instructions (alongside the connection-mode statement), an authoring_guidance field on the four pre-write reads (get_case_overview, get_argument_tree, suggest_placement, get_draft_status — the looped reads get_element and find_elements stay lean deliberately, so the guidance is not noise a model learns to skip), and the operation schema's text description, which is in context at the exact moment a claim's words are generated. A test resolves every id the doctrine names through the server's own sccg://guideline/<id> resource, so a condensation naming a guideline the catalog no longer has fails rather than quietly lying, and a second holds every rendered line to the published short_rule it must quote; the stdio smoke test proves the doctrine crosses the shipped binary's transport in both consent modes (the doctrine is the public house rules, not case content). The prompt-quoted sets also widened — the SU, LF and RD families were previously quoted in no prompt at all. |
| AF-AI-024 |
Mechanical SCCG checks bound to the published catalog |
supported |
src/core/sccg/staged_checks.cpp, src/app/areas/staged_finding_text.cpp |
tests/test_sccg_staged_checks.cpp, tests/test_staged_finding_text.cpp |
Phase 2 of docs/features/mcp-authoring-quality-plan.md. Fifteen checks, up from four: all five of SCCG's published deterministic pre-checks now run and are reported by their registry id (check-explicit-strategy, check-evidence-trace, check-evidence-citation-precision, check-evidence-control-attributes, check-evidence-state-fixed), plus EV.1 (unsupported claim not marked undeveloped), AR.2 (a decomposition whose sub-claims carry no reasoning step -- the catalog's check-explicit-strategy, advisory because GSN permits goal-to-goal support), AR.1 (solution with children; strategy developing into nothing; support cycles), CL.5 (unbounded qualifier), and batch 1 of the widening — CL.2 (two properties joined by a conjunction), CL.6 (lifecycle steps chained in one claim), RD.1 (reasoning smuggled into claim text), RD.4 (promotional language), EV.7 (evidence with no owner, version, date or status), EV.8 (mutable source cited with nothing fixing its state), LF.3 (absence of discovered evidence offered as support), and — new with SCCG 0.7.0 — CL.4 (a qualifier or hedge two reviewers can read differently) and CL.3 (a claim past the published word count, or a topic list written as prose). Every check binds to the catalog's suggested_checks id, carried on the finding and on the MCP wire. The word lists and thresholds are now read from the catalog rather than held here: SCCG 0.7.0 publishes tool.markers for 27 guidelines with three effects — candidate raises a signal, suppress cancels one (CL.5 no longer fires on a claim that states its bound), and expected inverts it, its absence being the signal — and tool.thresholds for CL.3 and LF.6. Two tools matching different words are not running the same check, and the hand-derived lists this replaced were narrower than the guideline in ways nothing recorded. ImplementedCheckIds() is now conditional on the catalog: a lexical check whose word list could not be loaded is not named, because an agent reading an empty findings array must be able to tell "found nothing" from "never ran". Each lexical check must fire on its guideline's own bad example and stay silent on the good one, read from the catalog at test time. A drift test holds each embedded statement to be a prefix of the catalog's, over a model built to trip every check the tool says it implements — it reached seven of thirteen before, so six quotes were never compared and one of them was a paraphrase. All lexical findings are advisory — the reviewer judges the words, the check only points — and the panels translate them by check id (the structured-findings seam), so the widening reaches the human review surfaces and MCP clients from one implementation. Partial catalog coverage: the remaining candidates are CL.1, EV.3, LF.1's textual half, LF.6, AR.8 and SU.2. CL.3 and CL.4 have left this list because 0.7.0 published the threshold and the word list they were waiting for. LF.6 has its threshold now and is still deferred for a different reason: the check needs a definition of "a figure" in claim text that a significant-digit count alone does not give, and a false precision finding against a correctly stated number is the kind that teaches a reviewer to stop reading. CL.6 fires on two lifecycle verbs joined by a conjunction where SCCG's published condition is two anywhere in the claim — a narrowing calibration SCCG permits, kept because it is what lets the finding quote the chain it objected to. tool.repair, published for all 48 guidelines, is read but not yet used to seed proposals. |
| AF-AI-025 |
Pre-flight rehearsal and submit-time refusal of problem findings |
supported |
src/agent/draft_operations.cpp, src/mcp/tools.cpp, src/mcp/session.cpp, src/core/drafts/draft_workspace_store.cpp, src/core/drafts/draft_persistence.cpp, src/ui/panels/draft_changes_panel.cpp |
tests/test_agent_request_handler.cpp, tests/test_mcp_modes.cpp, tests/test_draft_workspace.cpp |
Phase 3 of docs/features/mcp-authoring-quality-plan.md — where the guidance gains teeth. check_operations rehearses operations on a copy of the integrated draft: the same validation and findings stage_operations would return, with nothing stored, no revision moved, no element ids allocated, and nothing drawn on the user's canvas — an agent iterates privately until its work is clean, then stages once. It is the one draft-vocabulary tool that works offline, because a rehearsal against the accepted copy is a read; the byte-identical sweep covers it. Submit-time refusal: submit_change_group refuses while problem-severity findings (SCCG Problem checks and GSN well-formedness findings) stand against the group, naming each one in problem_findings. The gate sits at submit, not at staging — staging is deliberately incremental and every intermediate shape is legitimately unfinished, but submit is the author declaring itself done. The explicit escape is acknowledge_findings: true, which submits anyway and records the acknowledged findings on the group — persisted, surviving restart, shown to the reviewer on the Draft Changes row, and cleared when the group's operations change, because the waved-through findings described a different shape. Advisory findings never refuse anything, deliberately: gating on them would train agents to acknowledge reflexively, which would spend the gate. Promotion authority is untouched — the human reviewer's decision is not gated by anything here. |
| AF-AI-026 |
Offline SCCG review evaluation harness |
supported |
src/eval/sccg_review_eval.cpp, src/eval/sccg_review_eval_options.cpp, src/review/sccg/sccg_review_preparation.cpp, src/review/sccg/sccg_review_passes.cpp |
tests/test_sccg_review_preparation.cpp, tests/test_sccg_review_passes.cpp, tests/test_sccg_review_eval_options.cpp |
af-sccg-review-eval, a headless binary that runs the SCCG review method over a project from the command line: the same profile selection, data packages, pre-checks, request and response validator as the in-app review, with no window. It writes one JSON record per run — SCCG version, profile, guidelines carried, packages supplied and the availability state of each one absent, pre-check verdicts, model, prompt hash, elapsed time, and every finding with its cited guideline — which is the review record the SCCG method says a tool should retain. --runs repeats a review so a guideline that fires intermittently can be told from one that does not fire; --model varies the model without touching saved settings; --dry-run assembles and records the request without calling a provider, so profile selection, package availability and prompt content can be checked at no cost. It exists because the review method was previously reachable only by rendering a frame, which makes guideline coverage over a whole argument impossible to measure. It is not a gate: it calls a paid external provider and its output depends on a model, so it is never a CTest — review::PrepareSccgReview is what the tests cover, and this turns the prepared request into a recorded run. Review passes are sent concurrently and merged exactly as the application does; each record lists every pass with its guidelines, prompt hash and size, latency, outcome and raw response, plus any finding discarded for being cited outside its pass. --single-request sends the whole profile as one request instead, so the two can be compared on the same material; the application never does. Cost controls for sweeps (none of which the application uses): --service-tier flex sends OpenAI's slower tier at batch prices with a 900 s timeout; a refusal made before any work -- a rate limit, flex having no capacity, an overloaded server (HTTP 500/502/503) -- is retried with exponential backoff and the attempts are recorded, while a timeout (possibly billed) and an exhausted account are not. Every run record carries token usage per pass and per run, including cached input and reasoning, and the tier that actually served each request, so the cost of a sweep is in its records rather than only on the provider's invoice. --first-run n numbers an invocation's runs from n, so a three-run sweep can be extended to five without paying for the first three again; a dry run from both builds confirms the requests are byte-identical before runs from two invocations are counted together, and a consensus record names the first_run and last_run it covers. The command line is tested: every number is refused when malformed or out of range -- --timeout 0 used to switch the request deadline off, and --runs 4294967297 read as 1 -- and a --first-run/--runs pair whose last run number would not fit an int is refused rather than running nothing. A record that cannot be written counts as a failure, and an output directory that cannot be created stops the harness before any paid call. Every per-run finding carries the model's confidence, not only the consensus. user_prompt_bytes is the size of the text user_prompt_sha256 hashes, separators included; the passes' own prompts sum to passes_prompt_bytes. A run's usage states passes_reporting and whether it is complete, since a pass that reported none leaves the total short. tool_build is stamped at build time rather than configure time, so a commit in the same build tree no longer leaves every record naming the previous build. The retry policy itself has no deterministic test yet. |
| AF-AI-027 |
Consensus review over repeated runs |
supported |
src/review/sccg/sccg_review_consensus.cpp, src/eval/sccg_review_eval.cpp |
tests/test_sccg_review_consensus.cpp |
Runs the same unchanged review request k times and groups what comes back by the guideline it cites, reporting each finding with the number of runs that produced it, whether they were unanimous, every run's wording, and any deterministic pre-check that independently reached the same guideline. A floor (--consensus m) separates well-supported findings from weak ones; findings below it are kept and reported separately rather than dropped, because a finding one run of three raised is weak evidence and not none. Counting rules that matter: one vote per guideline per run (a run that raised a guideline three times has found one thing worth saying three times), the same guideline on different elements is different findings, and agreement is counted against the runs that succeeded — 2 of 2 when the third run errored is not 3 of 3. This does not make review deterministic and must not be described as doing so; it makes non-determinism visible. It exists because the provider offers no way to make a review repeatable: measured directly, gpt-5.6-sol rejects temperature ("not supported with this model") and seed ("unknown parameter"), gpt-5.5 rejects temperature too, and only gpt-5.4 accepts it — so the newer the reasoning model, the less sampling control there is. Currently exposed through the evaluation harness; the in-app review still runs once. A run in which any review pass failed counts as a failed run, not as one that cited nothing: its missing pass never asked about the guidelines it owns, and counting them as "not cited" would put a false disagreement into the consensus. |
| AF-AI-028 |
Sampling controls in AI settings |
supported |
src/ai/ai_types.h, src/ai/openai_provider.cpp, src/eval/sccg_review_eval.cpp |
tests/test_openai_provider.cpp |
temperature and seed are carried as optional settings and sent only when set, because "whatever the provider defaults to" is a real configuration and a model that rejects the parameter must still be reachable — the current reasoning models reject both, so a settings type that always sent a number could not talk to them. An unconfigured install therefore behaves exactly as it did before these existed. --temperature / --seed on the evaluation harness, and the model block of every run record states what was actually used, so a run's repeatability is a fact in the record rather than something the reader has to remember about the command line. |
| AF-AI-029 |
Prompt caching for SCCG reviews |
supported |
src/review/sccg/sccg_review.cpp, src/ai/openai_provider.cpp, src/app/controllers/ai_review_controller.cpp, src/eval/sccg_review_eval.cpp |
tests/test_sccg_review_preparation.cpp, tests/test_openai_provider.cpp, tests/test_ai_review_controller.cpp |
A review request is built as three segments, most shared first: what every review sends (instructions and the response contract, ~2.5k tokens), what every review of one profile and pass sends (the pass sentence, the profile and its rules, ~6-7k), and the element's own data (~1.6k). The text is unchanged -- only the order, with the response contract moved ahead of the element data -- and a test holds that nothing of an element reaches the first two segments. The application places explicit cache breakpoints after the first two and none after the element data, in explicit mode, so it never pays the 1.25x cache-write price for data a review of one element will not read back; prompt_cache_key names SCCG version, profile and pass so requests that share rules reach the same cache. Measured on gpt-5.6-sol (2026-09-11): a first claim reviewed wrote ~9-10k tokens per pass to the cache, and a different claim straight after read ~9k of its ~10.5k input tokens per pass from it, at 0.1x the input price. Requests are no slower: under the same provider load the cached requests were as fast or faster than uncached ones. What caching cannot reduce is output: a claim-review pass produces ~2k output tokens, about 70% of them hidden reasoning, and output is now most of the bill. The evaluation harness also caches each element's data when it reviews an element more than once (--runs > 1), since every later run repeats it, and --no-prompt-cache sends the old single string for comparison. Correction: that single string, and a prompt edited in the debug panel, were described as uncached but were not: with no breakpoints OpenAI places an implicit one at the end of the message. Both now opt out explicitly -- explicit mode with no breakpoints, which OpenAI documents as using no cache and writing none. A billed response whose output is unusable keeps its usage, and a usage block without both token counts is no longer reported as measured usage. |