Async: what runs in parallel¶
ZeroAI is single-process, single-thread, asynchronous. One event loop (uvicorn + asyncio) serves
every request. Nothing blocks it: network, database and model calls are all awaited, so while one
request waits for a feed or for the model, the loop serves another.
The one-paragraph version
Inside one briefing the steps run one after another, except fetching the news feeds, which runs in parallel. Across briefings everything overlaps, except the model call, which is deliberately one at a time (a semaphore). Closing the tab cancels whatever is running.
Where time goes in one briefing¶
gantt
title One briefing (illustrative seconds)
dateFormat HH:mm:ss
axisFormat %S s
section Feeds (in parallel)
Yahoo Finance :a1, 00:00:00, 2s
Google News :a2, 00:00:00, 3s
Nasdaq (slow) :a3, 00:00:00, 3s
section Model (one at a time)
Wait for the gate :crit, b1, 00:00:03, 1s
Reasoning + answer :b2, 00:00:04, 5s
section After
Store usage row :c1, 00:00:09, 1s
The three feeds overlap, so fetching costs as long as the slowest feed, not the sum. Then the request takes its turn at the model.
Two briefings at the same time¶
gantt
title Two requests overlap, but the model runs one at a time
dateFormat HH:mm:ss
axisFormat %S s
section Request A
Fetch feeds (parallel) :a1, 00:00:00, 3s
Model :a2, 00:00:03, 5s
section Request B
Fetch feeds (parallel) :b1, 00:00:01, 3s
Waiting for the model :crit, b2, 00:00:04, 4s
Model :b3, 00:00:08, 5s
B fetches its news while A is still working, so by the time the model is free B is ready. Only the
model call queues. The UI shows "Waiting for the model" while it does.
What is parallel, serial, or limited¶
| Step | How it runs | Why |
|---|---|---|
| Fetch each news source | In parallel: asyncio.gather over all enabled sources, one shared httpx2.AsyncClient |
I/O bound; total time is the slowest source, and one failure never blocks the others |
| Retry of a failed feed | Sequential per source, with asyncio.sleep(0.3) between the two tries |
Only timeouts, connect errors and 5xx are retried, once |
| Parse feeds, merge, de-duplicate | Inline on the loop, one after the other | A few milliseconds of CPU; not worth a thread |
| Model call | One at a time: LLMGate = asyncio.Semaphore(ZEROAI_LLM_MAX_CONCURRENCY) |
Protects a local model and Ollama's free cloud tier from being flooded |
| Reasoning and answer tokens | Streamed as they arrive; each becomes an SSE event | The user watches it live instead of waiting for the whole answer |
| Database reads and writes | awaited on aiosqlite (SQLite runs on a helper thread, off the event loop) |
The loop never waits on disk |
| Recording usage | After done, in its own session |
The request's session is closed by then |
| Different users' requests | All overlap, apart from the model gate | One loop serves them all |
| The browser | EventSource for the stream, React Query for REST, timers for the clock |
The page stays responsive while the model works |
One request, step by step¶
flowchart TD
R([GET /summary/stream]) --> D["Dependencies, awaited in turn:<br/>DB session, model config, enabled sources"]
D --> F["asyncio.gather<br/>fetch every source concurrently"]
F --> N[/"event news"/]
N --> G{"LLMGate<br/>free slot?"}
G -->|"no"| W[/"event status: Waiting for the model"/] --> G
G -->|yes| M["PydanticAI run_stream_events<br/>(runs the agent in a background task)"]
M -->|"reasoning tokens"| T[/"event thinking"/]
M -->|"answer fragments"| S[/"event summary"/]
M -->|"result + token usage"| U["Release the gate,<br/>write the usage row"]
U --> DN[/"event done"/]
The route is an async generator: each yield hands one SSE event to Starlette, which sends it to the
browser immediately. The generator only advances when the client is ready, which is natural backpressure.
The model call, in more detail¶
run_stream_events starts the agent loop in a background task and passes events to our generator
through an in-memory channel. Our code consumes them one at a time:
- a reasoning delta becomes a
thinkingevent (dropped on the server if Thinking mode is off); - a fragment of the structured answer is appended to a buffer, parsed as partial JSON, and sent as
a
summaryevent when it changed; - the final result carries the validated answer and the token usage.
If the model returns nothing usable, PydanticAI asks again (up to 3 retries) and our partial-JSON buffer is reset for the new attempt.
Cancellation¶
sequenceDiagram
participant B as Browser
participant S as Starlette
participant G as stream_summary
participant L as LLMGate
participant M as Model request
B->>S: stream open
S->>G: iterate events
G->>L: acquire (held)
G->>M: streaming request
B--xS: tab closed / reload / Stop
S->>G: cancel (CancelledError at the next await)
G->>M: request dropped
G->>L: release (finally)
Note over G: logs client_disconnected<br/>no usage row is written
Cancellation is cooperative: it lands at the next await, which for a model that streams tokens is
milliseconds away. A request cancelled while still queued never reaches the model and gives its place
back. This is covered by tests that cancel mid-call and while queued.
Shared resources, created once¶
| Resource | Created | Why shared |
|---|---|---|
httpx2.AsyncClient |
at startup (lifespan) | Connection pooling and keep-alive across all feed fetches |
AsyncEngine (SQLite) |
at startup | One pool; sessions are cheap and per request |
LLMGate |
at startup | The semaphore must be one object for the limit to mean anything |
What could go wrong, and why it does not¶
| Risk | Answer |
|---|---|
| A slow feed stalls everything | Per-attempt timeout (8 s), one retry, and the other feeds are unaffected |
| A CPU-heavy step freezes the loop | There is none: parsing a feed or a partial answer is microseconds to milliseconds |
| Many tabs pile up model calls | They queue behind the gate, and abandoned ones are cancelled |
| Lost writes under concurrency | SQLite serialises writers; our writes are tiny and awaited |
| Work continues after the user left | Disconnect cancels the task, which drops the model request |
Not parallel on purpose
More parallelism is not always better here. Raising ZEROAI_LLM_MAX_CONCURRENCY only helps for a
hosted API or a GPU that serves several requests at once; on Ollama's free cloud tier, parallel
requests were measured to queue on the other side (every run took about 380 s).