mirror of
https://github.com/fawney19/Aether.git
synced 2026-10-04 00:17:45 +08:00
Raw chain-of-thought was written to both `content` (`reasoning_text`) and `summary` (`summary_text`), and the stream emitter sent the same delta on `response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`. Clients that render both channels therefore printed every thinking chunk twice — most visibly the Codex CLI, whose thinking panel repeated itself. OpenAI keeps the two channels distinct: `content` carries the raw CoT while `summary` is the summarised view. Emit the thinking on `content` only: - `openai_responses_reasoning_text_fields` becomes `openai_responses_reasoning_text_parts`, returning just the `content` array; reasoning items keep `summary: []` (or a provider-supplied summary). - The Responses stream emitter emits `response.reasoning_text.delta` / `.done` and no longer mirrors them onto the summary events. The reasoning `output_item.added` no longer announces a `reasoning_summary_part`. - The provider-state reasoning reader accepts `content` (`reasoning_text`) first and falls back to `summary`, so it also understands items produced by older Aether versions; its state field is renamed accordingly. - Non-streaming builders (Chat -> Responses, manual Responses response, Grok gateway) place the thinking on `content` and leave `summary` empty. Tests cover the raw thinking appearing exactly once in the emitted stream.