Thinking mode¶
Some models "think" before answering. ZeroAI can show that reasoning live in a panel above the briefing. The Thinking mode switch (Settings, Ollama only) controls it.
| Switch | What ZeroAI does |
|---|---|
| ON (default) | Asks for reasoning at the chosen effort, streams it to the UI as thinking events, collapses it when the answer starts. |
| OFF | Does not ask for reasoning and never shows any. Faster and more reliable. |
Thinking effort¶
With Thinking mode on, choose how hard the model thinks: Low, Medium or High (default).
Measured on gpt-oss:120b-cloud, a full NVDA briefing from live feeds:
| Effort | Reasoning shown | End to end | What it is for |
|---|---|---|---|
| Low | about 260 characters | about 4 s | The quickest answer; the prompt also asks for brevity |
| Medium | about 800 characters | about 4 s | A short paragraph of reasoning |
| High (default) | about 4,500 characters | about 6 s | Thorough reasoning you can read |
Which models honour it
gpt-oss follows all three levels. qwen3 only has on or off, so any effort means "think".
The effort selector is hidden when Thinking mode is off, and for OpenAI.
What each model actually does¶
flowchart TD
A{Thinking mode} -->|ON| B["reasoning_effort = low, medium or high<br/>(brevity prompt only for low)"]
A -->|OFF| C{Can the model<br/>disable reasoning?}
C -->|"yes: qwen3"| D["reasoning_effort = none<br/>no reasoning is produced"]
C -->|"no: gpt-oss"| E["reasoning_effort = low<br/>reasoning produced but<br/>hidden by the server"]
B --> F[thinking events reach the UI]
D --> G[no thinking events]
E --> G
| Model | ON | OFF |
|---|---|---|
gpt-oss:* (cloud) |
Reasoning at the chosen effort, shown | Lowest effort (cannot be disabled), hidden |
qwen3 (local) |
Reasoning shown; ~30 s; may return nothing | No reasoning; much faster, reliable |
| Models without reasoning | Nothing to show | Same |
| OpenAI | Switch is hidden; no reasoning is requested | n/a |
If you see \"did not return a usable answer\"
Turn Thinking mode off (or switch model). A reasoning model sometimes spends its whole turn thinking and returns nothing; ZeroAI retries up to 3 times, but each attempt costs a full thinking pass.
Why 'off' hides instead of only disabling
Ollama's OpenAI-compatible endpoint maps reasoning_effort: none to "do not think", but
gpt-oss ignores it and always reasons. So the server also drops thinking events when the
switch is off, which makes the switch behave the same for every model.
Where it is stored¶
LLMConfig.thinking (default on) and LLMConfig.thinking_effort (default high) in the database. It is part of PUT /api/v1/settings/llm; omit it
to keep the stored value. Databases from before this setting existed are upgraded in place.