Where coverage is partial — particularly OpenAI’s integrity-checkpoint coverage — Mnemom does not claim parity. The v1 commitment is honest per-provider differentiation, not uniform coverage. Marketing materials and the trust center reflect this.
Supported models
Models the v1 promise applies to. Listed in the gateway’s/models.json registry; routed end-to-end through Safe House, AIP, CLPI, and DLP per the matrix below.
Anthropic
Claude Fable 5.1
Claude Opus 5.5
Claude Opus 5
Claude Sonnet 5.5
Claude Sonnet 5
Claude Opus 4.8
Claude Fable 5
Claude Opus 4.7
Claude Sonnet 4.6
Claude Haiku 4.5
OpenAI
GPT-5
GPT-5 Codex
o3
o3-mini
GPT-5.6 Sol
GPT-5.6 Terra
GPT-5.6 Luna
GPT-6 Astra
GPT-6.1 Sol
GPT-6 Sol
GPT-6 Luna
GPT-5.5
Gemini
Gemini 2.5 Pro
Gemini 2.5 Flash
Gemini 3.8 Flash
Gemini 3.5 Flash-Lite
model: on a Mnemom API call):
Feature coverage matrix
Each cell describes the v1 commitment level. Symbols:- ✓ Fully supported. Tested in CI; Safe House features work identically to the Anthropic baseline.
- ⚠ Partial — Provider exposes the feature, but with a documented limitation. See the supporting note.
- N/A — Provider does not expose this capability. Not a Mnemom limitation.
[1] Anthropic — full extended thinking
Anthropic models expose full extended thinking blocks through the response API. AIP reads completed thinking blocks post-response (before delivery inenforce, which adds latency; post-delivery in observe/nudge), and the verifier (Claude Haiku 4.5) has full chain-of-thought visibility. This is the most complete AIP coverage Mnemom offers — and the baseline against which other providers are compared.
If the client’s thinking settings return thinking blocks without readable text, AIP grades the turn’s visible reply and tool calls instead of skipping the check. See Integrity Checkpoints.
[2] OpenAI — no thinking-trace inspection through the gateway today
OpenAI’s reasoning summaries are a Responses API feature (reasoning.summary), not something the Chat Completions API returns. The gateway’s default /openai door proxies Chat Completions — the shape OpenAI clients send by default — and that shape never carries reasoning content back, for any OpenAI model, reasoning or not. AIP’s thinking-extraction step returns empty by construction for every /openai/* request today, including through the /openai/v1/responses passthrough path. There is no per-model split here: it is uniform across o3, o3-mini, gpt-5, gpt-5-codex, gpt-5.6-*, gpt-6-astra, gpt-6.1-sol, gpt-6-sol, gpt-6-luna, and gpt-5.5.
What this means in practice:
- Boundary violations that surface in the model’s final response are still caught equally well on OpenAI — Safe House and CLPI don’t depend on the thinking trace.
- Boundary violations that would only have surfaced in the model’s hidden reasoning are not caught on OpenAI — AIP has nothing to analyze, so a checkpoint below the minimum-evidence threshold is emitted as a synthetic
clearwithout an analysis-LLM call, not a partial-confidence verdict. - If your application relies on AIP catching boundary violations in the model’s reasoning specifically, use Anthropic or Gemini for that workload. OpenAI models remain fully supported for inference, Safe House, and CLPI — just not for thinking-trace inspection.
[3] Gemini — full thoughts exposure
Gemini models expose a thoughts field on response candidates. Coverage is uniform across 2.5 Pro, 2.5 Flash, 3.8 Flash, and 3.5 Flash-Lite — all four ship Google’s thinking capability (3.8 Flash defaults to medium effort, 3.5 Flash-Lite defaults to minimal, both adjustable) — and AIP reads thoughts through the gateway’s response normalizer and treats it as equivalent to Anthropic extended thinking.
[4] Anthropic — explicit cache_control
Anthropic supports explicit cache_control block markers — customers control which prompt segments are cached. Mnemom passes cache_control through transparently. Safe House still evaluates the full request (it does not assume cached prefixes are safe just because they were previously seen). Cache hits do not bypass any checkpoint.
[5] OpenAI — automatic caching, no customer control
OpenAI’s prompt caching is automatic — the API decides what to cache based on request shape. Customers cannot reason about cache hit rates the way they can on Anthropic. Mnemom passes requests through unchanged; cache decisions are OpenAI’s. Safe House dispatch remains idempotent across cache hits and misses — the same prompt routed through Mnemom twice produces the same verdict regardless of whether OpenAI cached it.[6] Gemini — separate CachedContent API
Gemini exposes prompt caching as a separate CachedContent resource (explicit cache lifecycle, named caches with TTL). The gateway today does not surface or use this API; requests are sent without referencing cached content. Customers using Gemini’s cache outside Mnemom will see lower latency than they see through Mnemom.
AIP coverage by provider — the headline commitment
Of all per-provider gaps, thinking-trace coverage is the load-bearing one:
If your application relies on AIP catching boundary violations in the model’s reasoning — particularly violations that would not surface in the final response — choose Anthropic or Gemini. OpenAI models are fully supported for inference, Safe House screening, and CLPI tool-use governance; they are not the right choice when AIP thinking-trace analysis is the load-bearing safety layer.
Latency expectations
In aggregate:- Safe House dispatch adds ~15 ms P50 / ~60 ms P95 across all providers (the dispatch path is provider-agnostic).
- AIP analysis runs post-response (before delivery in
enforce, which adds the latency below; post-delivery inobserve/nudge), only when a thinking trace was extracted. Cost varies by upstream-response token volume — Anthropic Opus with full extended thinking emits the largest traces and therefore the longest AIP analysis tails (P95 up to 2.5 seconds). Gemini 2.5 Flash emits thinner traces and completes AIP analysis in P50 ≤ 800 ms. OpenAI requests skip the analysis-LLM call entirely (no trace to analyze) and add no AIP latency.
enforce, AIP runs after the upstream response completes but before the gateway delivers it to the customer, so a violating response is gated before delivery (adding latency); this holds on every provider. In observe and nudge, the response is delivered first and the verdict is recorded post-hoc.
How we test against each provider
Mnemom’s gateway adapter — the code that parses each provider’s response format, extracts thinking blocks, and routes tool calls — is tested in two layers:
All three providers (Anthropic, OpenAI, Gemini) run both layers. Live tests fire nightly against real upstream APIs and detect breaking changes within 24 hours.
OpenAI reasoning content, streaming or not: OpenAI’s Chat Completions API — streaming or non-streaming — does not return reasoning content for any model. Server-side reasoning token usage is reported in
usage.completion_tokens_details.reasoning_tokens, but the reasoning text itself never reaches the response body Mnemom’s gateway proxies. This is upstream API behavior, not a Mnemom limitation, and it is why OpenAI has no thinking-trace coverage today regardless of streaming mode — see [2] above.Deprecation policy
Models the gateway routes are classified into two tiers:Supported
Listed in the supported models section above. v1 promise applies. Safe House, AIP, CLPI, DLP all work to their per-provider commitment. Tested in CI. Deprecation requires 90 days’ notice via this page and the changelog.
Passthrough
Routed by the gateway but not in the supported tier. Inference works; Safe House features are best-effort. Not tested in CI. No deprecation notice — model availability tracks the provider’s lifecycle.
Passthrough-tier models that are removed from the upstream provider also disappear from the gateway’s
/models.json registry; we do not maintain shims.
Out of scope
Provider expansion beyond Anthropic + OpenAI + Gemini is not in scope for v1. Cohere, Mistral, Together, Groq, and other providers are not supported. Adding a provider is a multi-quarter effort (Safe House dispatch, AIP adapter, CLPI policy schema mapping, test coverage, docs) — tracked separately. BYOK (bring-your-own-key) for upstream providers is also out of v1. v1 ships with Mnemom-only key custody.Related
- Integrity Checkpoints — the AIP analysis machinery
- Safe House — screening layer for everything reaching the agent: the inbound message at request arrival, plus each tool result turn-internally, inline, before the request is forwarded
- CLPI — Continuous Local Policy Interpretation for tool use
- Webhook contract — event delivery for operator surfaces (provider-agnostic)