Usage records show a reasoning badge next to the model name for OpenAI and
Claude requests, but Gemini requests never got one. The extraction only read
the OpenAI/Claude shapes (`reasoning_effort`, `reasoning.effort`,
`output_config.effort`), while Gemini states its reasoning depth inside
`generationConfig.thinkingConfig` — so nothing was written to the usage
metadata and the list and detail views had no badge to render.
Read the Gemini shape too, as a fallback after the existing three so the
OpenAI and Claude paths are untouched:
- `thinkingLevel` / `thinking_level` wins when present, trimmed and
lowercased, with the protobuf enum prefix stripped so `THINKING_LEVEL_HIGH`
resolves like `high`.
- Otherwise `thinkingBudget` / `thinking_budget` goes through the existing
shared budget ladder, yielding the same `low|medium|high|xhigh` vocabulary
the badge already understands.
- Both camelCase and snake_case spellings are read, so a captured client body
and a converted provider body resolve to the same label.
- `includeThoughts` alone is a visibility flag, not a depth, and produces no
badge.
Two cases are handled explicitly rather than through the shared ladder:
- `thinkingBudget: 0` disables reasoning outright. The shared ladder maps
`0..=1664` to `low`, which would report an explicitly disabled request as a
shallow one, so it reports `none` instead.
- `THINKING_LEVEL_UNSPECIFIED` is the enum's "no explicit level" member, not a
depth; it is rejected rather than surfaced as an `unspecified` badge.
The frontend needs no change: `UsageModelDisplay` already renders the badge
whenever the fields are present, and keeps the `high -> xhigh` mapping format
when the requested and upstream efforts differ.
Consolidate subscription usage policy enforcement, privacy-safe persistence, and gateway security hardening into one reviewable change.
Includes bounded HTTP and execution envelopes, header and protocol guards, DNS and relay validation, authentication and secret projection hardening, secure backup/install paths, and regression coverage.
Retry pre-response transport failures across candidates with an explicit stop policy, and propagate end-to-end timing into usage records and UI diagnostics.
Remove legacy body, import, cookie, PII, and tunnel replay caps while preserving optional operator-configured gateway limits.
Shard and singleflight hot-path caches, batch and prioritize candidate and usage lifecycle persistence, and extend database and pressure-test instrumentation for 20k concurrent streams.