Atomic Behaviors Catalog
Rather than duplicating monolithic rule sets for every model vendor, the validation battery defines 21 atomic behavior profiles. Model family recipes compose these granular profiles as pure data. The alternative — maintaining one giant bespoke rule blob per vendor — is how codebases end up with 40 copy-pasted files that silently drift out of sync the first time someone fixes a bug in one and forgets the other 39.
The Doctrine: Stated vs. Measured Reality
Every rule in this catalog was originally derived from vendor documentation and assumed to represent a hard wire constraint. A live audit then dispatched every rule against its own vendor's native API — not through a gateway or proxy — and measured what the vendor actually did. The result: 16 of 17 rules blocked turn state the vendor accepts. A gateway's automatic repair and a vendor's actual tolerance look identical from the outside; only native testing reveals what the model stack genuinely rejects.
The code reflects this ground truth: rules are advisory by default (severity: 'advisory'). They record findings for observability rather than gating dispatches. You opt into gating per rule with severity: 'blocking'. The documentation-derived rules remain in the catalog because they record which primitive breaks a vendor's stated contract — but each profile below documents what was actually measured.
The 21 Atomic Profiles
1. 'permissive'
An empty rule set (rules: [], permissive: true). This captures the deliberate, documented design choice by xAI Grok to impose zero role-order limitations on conversation history. In the live audit, this profile served as the control: the guard remained silent, and Grok's native API accepted the dispatches as designed.
2. 'openai_shape_baseline'
The baseline tool-calling guard for OpenAI-compatible conversations. Enforces that a Message primitive may not immediately follow a ToolCall primitive, reflecting the ADK's internal convention where tool execution results are stored directly on ToolCall rather than on correlated messages.
Audit reality: Measured against native APIs (Nova, GPT-4 legacy) and confirmed unfounded in practice — both native APIs accepted requests where a message followed a tool call without issue. The profile is advisory by default, recording when turn state diverges from ADK representation conventions without rejecting dispatches the model accepts.
3. 'strict_alternation'
Enforces that conversation messages alternate strictly between user and assistant roles (roles: ['user', 'assistant'], mode: 'strict').
Audit reality: Measured against Nova (with alternationPolicy: 'reject') and Gemma 4 on their native APIs; both vendors accepted non-alternating consecutive turns without rejection. An earlier audit through a gateway had appeared to confirm the rule, but that was an artefact: the gateway's own enforceAlternation() middleware was merging consecutive same-role turns before Bedrock ever saw them. On the native API, the rule is unfounded in vendor enforcement and advisory by default.
Field note: the trailing-assistant turn that no one rejects, until one model just stops
A violation of this rule doesn't always look like a 400. ADK's own "Punching Above Its Weights" showcase agent synthesizes plan/nudge Thoughts and injects them into the model-facing prompt as assistant-role content — generated in response to the user's turn, so they land last. The prompt ends on an assistant turn. Gemma and DeepSeek shrug and keep generating past it. qwen3-coder-next's chat template treats a trailing assistant message as a completed turn and stops immediately — no error, no violation, just eval_count: 1, empty content, and a dispatch loop that spins hundreds of times waiting for an answer that will never come. The fix wasn't a repair rule; it was a per-model opt-out that keeps synthetic thoughts out of the model-facing prompt (they still render in the UI) for chat templates known to end a turn on a trailing assistant message. strict_alternation catches the well-behaved failure mode — roles actually out of order. It says nothing about a technically alternating sequence whose last entry is the wrong role for a specific model's template, and that gap doesn't produce a rejectable violation at all: it produces silence.
Field note: the same trailing-assistant turn, at production scale, across three vendors
This isn't a showcase-app curiosity. A real production CI code-review agent hit the exact same shape twice, independently, and had to fix it in production. In one incident, a corrective "nudge" message followed by a synthesized echo Thought left the wire request ending on an assistant turn. Nova/Bedrock didn't reject it: it returned a well-formed 200 with content: null and finish_reason: "stop", byte-identical across four straight retries, silently burning the whole retry budget on a request the model was never going to answer. Other vendors' translators reject this same shape outright as a hard 400 ("Requests ending with a model turn are not supported") — so the failure isn't a Nova defect, it's undefined behavior that different serving stacks resolve differently, some loudly and some not at all. The fix there was structural: append a minimal, inert trailing message after any synthesized assistant-role content so the wire array always ends on a non-assistant turn.
The second incident was a variant of the same root problem with sharper numbers: a model that had already finished its real work re-presented with an unchanged prompt returned empty generations 30/30 times in a live-replay measurement (a clean reproduction of a real production outage) — not because anything was malformed, but because asking a model to repeat itself when it has nothing left to say is, itself, an ordering failure of a subtler kind. Injecting one corrective turn and re-presenting the model's own prior output as a Thought rather than a Message dropped that failure rate from 30/30 to 8/30 in the same live measurement. This battery's strict_alternation profile is the validation-time shape of exactly this lesson: a request that alternates roles by the letter of the rule can still end on the wrong note in the eyes of the specific model reading it, and the only way anyone found that out was watching it fail in production, repeatedly, at scale, until it wasn't a mystery anymore.
4. 'single_tool_call_per_turn'
Enforces a cardinality cap of at most one ToolCall per assistant role group (maxPerGroup: 1), matching Meta Llama 3's documented non-parallel tool-calling constraint.
Audit reality: Documentation-derived and advisory by default. It captures Llama 3's stated architectural limit without blocking dispatches unless explicitly opted into severity: 'blocking'.
5. 'thinking_before_tool_use'
Enforces that within the active assistant turn (onlyLatestGroup: true), any Thought primitive must precede any ToolCall primitive, derived from Anthropic's extended-thinking documentation.
Audit reality: Measured directly against Claude Haiku 4.5 via the native Anthropic Messages API and confirmed unfounded — the native API accepted dispatches where tool calls preceded thoughts without error. The profile remains advisory by default, reporting deviations from Anthropic's documented sequence without rejecting dispatches the model accepts.
6. 'thought_signature_required'
Requires the first ToolCall in an assistant group to carry a thoughtSignature in its payload.
Audit reality: The ONE rule in this catalog confirmed by real vendor enforcement. When dispatched against Gemini 3's native API, an unsigned function call in replayed history was rejected with a hard HTTP 400 naming both the missing field and its position ("Function call is missing a thought_signature in functionCall parts ... position 2"). Sending the identical history with Google's documented sentinel returned 200.
Because measured enforcement exists, this profile keeps severity: 'blocking' explicitly while the rest of the catalog defaults to advisory. In action: 'mutate' mode, it repairs missing signatures without requiring the global allowMetadataFallbackRepair opt-in because it sets fallbackRepairAuthorized: true using Google's own published sentinel ('skip_thought_signature_validator'). See Gemini Sentinels.
7. 'thought_signature_advisory'
The advisory variant of 'thought_signature_required' (severity: 'advisory'). Evaluates the presence of thoughtSignature for Gemini 2.5 without blocking dispatch.
Audit reality: Measured against Gemini 2.5's native API and confirmed unfounded — Gemini 2.5 accepted function calls without thought signatures. The advisory posture matches reality: missing signatures are reported for observability without gating calls.
8. 'function_response_adjacency'
Enforces Gemini's documented tool-sequence adjacency rule: a Message may not immediately follow a ToolCall primitive before function execution flow concludes.
Audit reality: Measured directly on Gemini 2.5's native API and found unfounded — Gemini 2.5 accepted turns where a user message followed a tool call. In this ADK, tool results live on ToolCall itself, so an injected notice or subsequent user message following a tool call is normal turn assembly. Advisory by default; in action: 'mutate' mode, an opted-in blocking instance repairs via reorder-adjacent.
9. 'full_history_preservation'
Parameterized profile factory (full_history_preservation:<kind>). Stateful preservation check ensuring that historical primitive counts ('toolCall', 'thought', or 'message') never decrease across dispatch iterations.
Audit reality: In the native audit, this profile read JUSTIFIED through AWS Bedrock Converse and UNJUSTIFIED on Mantle — the difference was the gateway translation layer, not the underlying model. This result demonstrates why API surface scoping matters: a proxy may enforce history invariants that the native API ignores. Advisory by default; preservation rules are deliberately unrepairable in mutate mode because dropped context cannot be fabricated.
10. 'payload_field_preservation'
Parameterized profile factory (payload_field_preservation:<field>). Stateful check ensuring that specific opaque metadata fields (such as Anthropic thought signatures or GLM clear_thinking flags) remain identical across dispatch iterations.
Audit reality: Measured against GLM-4.7 and Codex native APIs and confirmed unfounded — both accepted dispatches where payload metadata had been modified or dropped. Advisory by default.
11. 'reasoning_pruned_after_latest_turn'
Implements Qwen 3's documented reasoning retention invariant: reasoning predating the latest non-tool-call user turn may be pruned, but all reasoning generated at or after that boundary must remain intact and stable.
Audit reality: Measured against Qwen 3's native API and confirmed unfounded — Qwen 3 accepted dispatches where reasoning retention diverged from the template specification. Advisory by default.
12. 'stale_thinking_advisory'
Implements Gemma 4 hygiene recommendations. Emits a non-blocking advisory if thinking content older than the latest user turn is resent in history.
Audit reality: Advisory by design (StaleContentAdvisoryRule); in the audit, both legs were accepted by Gemma 4.
Field note: this isn't theoretical — a shipped agent enforces it as a hard rule
The Gemma model card's guidance against resending prior-turn thinking isn't an edge case someone might hit. ADK's own "Punching Above Its Weights" showcase agent (Gemma-4 via LiteRT-LM) treats it as load-bearing: a prior turn's model-generated reasoning is stripped from every subsequent turn's prompt, full stop — no advisory, no opt-out, because the code comment citing it names the exact clause ("Gemma model card §3 — 'No Thinking Content in History'"). Only harness-authored synthetic thoughts (the plan, nudge corrections) survive replay, via an explicit allow-list, precisely because they aren't the model's own past chain-of-thought. This profile ships as advisory — informational, never blocking — because the vendor guidance itself is a recommendation, not a spec-enforced rejection. A caller who has actually watched a Gemma-family model degrade on stale thinking may reasonably decide this specific rule deserves to gate dispatch in their own pipeline rather than just warn.
Field note: found while debugging something else entirely
A real production CI code-review agent discovered this exact failure class sideways, while root-causing an unrelated empty-generation incident. Every LLM battery it used defaulted to replaying every self-authored plain-text thought from every prior turn, forever, unbounded — and none of its own harness-authored notices carried the metadata needed to opt out of that replay. The fix was to keep only the single most recent self-authored thought per turn instead of accumulating the whole history of them, explicitly called out in the commit as "the recommended posture against the unbounded-replay degradation some model families (e.g. Gemma) exhibit." Nobody set out to build that fix. It fell out of chasing a model that kept returning empty responses, and the actual root cause turned out to be a pile of the model's own stale reasoning it had never been told to stop re-reading. The advisory in this profile is the polite version of that lesson; the production incident is what happens when nobody's listening to it.
13. 'role_remap_split_tool_roles'
Requires every ToolCall to carry a producer-set role-remap marker confirming it has been rendered under Granite 3.x's distinct tool-call and tool-response wire roles.
Audit reality: Evaluated on Granite 3.x native API; both legs were accepted. The tag is a consumer-supplied payload field (payload.roleTag) that nothing in the ADK writes automatically. A blocking default would reject every tool call for consumers who have not hand-populated this field, so this profile defaults to advisory.
14. 'role_remap_inline_tool_call'
Requires every ToolCall to carry a producer-set role-remap marker confirming it has been rendered under Granite 4.x's inline tool-call, remapped-response wire roles.
Audit reality: Evaluated on Granite 4.x native API; both legs were accepted. Advisory by default for the same reasons as Granite 3.x.
15. 'harmony_commentary_channel'
Validates the OpenAI GPT-OSS Harmony format. Requires every function ToolCall to carry a payload.channel field, matching Harmony's commentary-channel tagging requirement.
Audit reality: Measured against GPT-OSS native API and confirmed unfounded — GPT-OSS accepted tool calls lacking the channel tag. Advisory by default. It provides 'commentary' as a fallback value so that mutate-mode repair works if opted into blocking.
16. 'converse_text_before_tool_use'
AWS Bedrock Converse hosting rule: within an assistant turn, all text Message primitives must precede any ToolCall primitives.
Audit reality: Deserves particular honesty. This rule was named after Bedrock Converse, was tested directly on Bedrock Converse, and Converse accepted BOTH orderings without complaint. The documented constraint does not exist on the native API. The profile is advisory by default.
17. 'tool_identity'
Requires every replayed ToolCall to name a tool that the current request actually declares.
Audit reality: Measured from live vendor behaviour, not read from documentation. This profile exists because of a root cause identified in production: Gemini matches functionResponse.name against the request's functionDeclarations. When a replayed tool call names a tool that resolves to nothing — such as when an opaque call ID is passed in place of the tool name, or when a tool offered on a previous turn is no longer offered this turn — Gemini returns an empty candidate (parts: [{text: ''}], finishReason: STOP, no candidatesTokenCount) a large fraction of the time.
A gateway forwards that response as an ordinary finish_reason: "stop" with content: null and no error, leaving the caller to record a successful HTTP 200 turn that produced nothing and spin in an empty generation loop. That silence is what makes it dangerous: there is no status code to catch and no error body to classify. Requires ctx.tools; skips silently when no tool registry is provided. Advisory by default, but a strong candidate for severity: 'blocking' on Gemini pipelines.
18. 'schema_integrity'
Requires every key in a tool schema's required list to exist in its properties.
Audit reality: Measured from live vendor behaviour, not read from documentation. This catches the other root cause of silent empty failures, and the more insidious one. A schema whose required array names a key absent from properties cannot be satisfied by any argument object. When presented with this unsatisfiable shape, Amazon Nova responds with a normal HTTP 200 that silently omits the field — 25 production responses were measured silently missing the field with zero errors at any layer.
The usual cause is a schema-sanitising pass that strips an annotation (like title) from properties without pruning the matching entry from required. Unlike turn-ordering rules, schema_integrity inspects declared tools in ctx.tools before a single token is generated. Advisory by default.
19. 'non_empty_turn'
Parameterized profile factory (non_empty_turn or non_empty_turn:terminal). Requires an assistant turn to carry either prose content or an adjacent tool call.
Audit reality: Measured from live vendor rejections. Two native APIs reject an assistant turn that carries neither content nor a tool call, in two different ways:
- Mistral rejects with an explicit HTTP 400:
"Assistant message must have either content or tool_calls, but not none." - Gemini rejects a request whose terminal
modelturn carries only athought: truepart withfinishReason: MALFORMED_RESPONSE(measured 4 of 4 times, against a normalSTOPwith generated text when the identical history ends on the user turn).
A Thought primitive alone does not count as content — that is precisely the shape Gemini refuses. The factory accepts onlyTerminal: true (for Gemini's terminal-position constraint) or false (for Mistral's history-wide check) and role (defaults to 'assistant'). Advisory by default.
20. 'tool_call_id_format'
Parameterized profile factory (tool_call_id_format:<maxLength>:<allowedPattern>, defaulting to 64 and '[A-Za-z0-9_-]'). Constrains ToolCall identifiers to provider-compatible lengths and character sets.
Audit reality: Measured from live vendor rejections. Two providers enforce hard identifier constraints that name neither the field nor the offending character, and both fail across every credential in a pool:
- OpenAI Codex returns HTTP 400 for any tool call ID longer than 64 characters. A production gateway translator triggered this when generating composite IDs embedding a UUID plus an iteration counter.
- Bedrock Converse rejects any
toolUseIdcontaining characters outside[A-Za-z0-9_-].
ADK's internal uuidv6 identifiers satisfy both constraints; this profile guards against consumer-supplied or cross-provider replayed IDs. Advisory by default.
21. 'tool_call_id_uniqueness'
Requires every ToolCall identifier to be unique across the dispatch timeline. This is a cross-entry set property, not a per-entry predicate: a finding names a collision group (three calls sharing one id is one defect, not three pairs), and the evaluator scans the whole timeline because the collision is definitionally cross-turn — any role- or group-scoped check would partition the colliding pair apart and see nothing wrong in either half.
Audit reality: MEASURED against xai.grok-4.3 on Bedrock Mantle. This upstream resets its tool-call counter per turn: five sequential tool-calling turns each returned call_0. Parallel calls within one response are numbered correctly (call_0..call_3), so the collision is strictly across turns — and the rule is deliberately silent on that correctly-numbered batch, which is what separates it from a rule that merely counts duplicates. Well-behaved upstreams already emit globally-unique ids, so the rule is a no-op for them and self-limits to actual collisions.
This profile is BLOCKING, unlike almost everything else in the catalog. An advisory finding can never be repaired here — repairViolations accepts blocking violations and middleware passes only the blocked set — so an advisory rule with a repair is a contradiction. Reusing an id corrupts result correlation and can make a later dispatch reject the completed call, so the collision is unsafe to assemble past.
It applies to the dispatch surface only (surface: 'dispatch'). The turn guard runs once pre-loop, before any dispatch has adopted a call, so the cross-turn collision it would repair is hydrated history rather than anything this turn produced — the dispatch guard sees the same state and repairs it in the right place. A consumer who wants turn-level detection sets surface: 'both' on their own instance of the standalone profile, at which point the turn guard reports and, in mutate mode, repairs the collision through the same group operation; under enforce it reports and aborts — ordinary blocking-rule behaviour under enforce, not something special to this rule or surface.
mutate mode rewrites tool-call ids in durable storage. The repair renames every member of a collision group through the context's atomic group operation, which persists the renamed calls. Anything a consumer has recorded against a call id — logs, traces, join tables, their own spool keys — can therefore point at an id that no longer exists. That is consented rather than imposed: action: 'mutate' is the switch, and the same class of act as the existing reorder repair, which rewrites createdAt on persisted primitives. See Operating Modes & Repair Strategies for the durable-id-rewrite consequence spelled out where a reader chooses between enforce and mutate.
A Known Gap: Stateful Primitive IDs Are Not an Ordering Concern (Yet)
The 21 profiles above validate sequence: what kind of primitive comes before, after, or adjacent to what. None of them validate identity continuity — whether a primitive that plays a stateful role across turns (a book-end plan, a nudge correction, any harness-synthesized control primitive a caller's own dispatch loop depends on being re-seeded turn after turn) keeps a consistent, non-colliding id as it's persisted and replayed.
A validation library that implies completeness it doesn't have is worse than one that admits a hole. We document this limitation plainly because it represents an unmodeled failure mode, not a solved problem.
Field note: an id collision can silently disable a book-end contract
This gap surfaced building ADK's own "Punching Above Its Weights" showcase agent, which opens every turn with a synthetic planning Thought and validates against that plan at turn close — a book-end contract. The in-turn id used to de-duplicate that thought while the turn is live (e.g. __plan-thought) is stable on purpose, so the live UI and the in-turn record map can find it. But that same stability becomes a liability at persistence time: if the thought were ever stored under that in-turn id instead of a freshly minted one, the next turn's history fetch would return a thought already carrying id __plan-thought — and the injection guard that decides whether to seed a fresh plan thought for the new turn (!seededIds.has(PLAN_THOUGHT_ID)) would see that id as already present and skip re-injection. The planner would silently stop running from turn 2 onward. No violation fires. No rule in this battery — not full_history_preservation, not reasoning_pruned_after_latest_turn — is shaped to catch it, because nothing was dropped or reordered; a primitive's identity just leaked across a boundary it was never supposed to cross. The fix was disciplined at the source (mint a fresh id at persist time, never reuse the in-turn key). The battery now has a repair strategy for the collision half of this class — renumber-colliding-ids renames a group of distinct primitives that wrongly share an id (see tool_call_id_uniqueness above) — but that repair does not cover this book-end case, where a single primitive's identity leaked across a boundary it was never supposed to cross. That continuity gap is still on the caller to defend.
Field note: the same shape, at higher stakes, in a real production panel
A stateful-id collision doesn't need a book-end contract to bite — it just needs two things that were supposed to stay distinct sharing an identity they shouldn't. A real production CI code-review agent measured exactly this: an orchestrator authored three genuinely distinct findings in one turn, and a downstream identity-collapse step folded all three under one reused id (the model, having been burned once by a rejected empty-string id, learned that a specific placeholder string was "accepted" and then reused that same placeholder for every finding that turn). Two of the three were silently discarded before they ever reached a human reviewer — a disposition gate built specifically to stop unbacked assertions instead dropped authored, backed ones, for a reason that had nothing to do with their content. The fix made identity issuance server-authoritative: an id is only honored as "already claimed" if the harness itself issued it and nothing earlier in the same batch already claimed it first, so a model reusing, guessing, or omitting an id always mints something fresh rather than colliding with someone else's. The failure direction was inverted on purpose — from "silently drop the duplicate" to "when in doubt, keep both." This battery's Known Gap is that same lesson at the primitive level: nothing here yet stops a caller's own stateful id from being replayed into a context it doesn't belong in, silently discarding whatever was supposed to live at that identity instead.
This is flagged here deliberately, not quietly worked around: it is a real, observed failure class that the current OrderingRule vocabulary does not model. If your own dispatch loop relies on a stateful control-plane primitive surviving replay by id, that continuity is on you to defend today — the battery does not yet have a rule shape for it.
See Also
- Validation Hub — Overview and model-to-recipe lookup table.
- Rule Types Reference — Full declarative schema for all twelve rule variants.
- Family Recipes Catalog — Matrix of all 38 pre-configured family recipes composing these behaviors.
- Writing a Profile — How to create new atomic profiles and recipes.
- Operating Modes & Repair Strategies — Configure enforce vs mutate and understand automated timeline repair.