NEVER TRAINED ON · PRIVATE BY DEFAULT · BUILT BY MOZILLA

GITHUB

Sources to answers

Trace a question from request to source-backed answer

Record a Tabstack Research request, preserve its cited pages, and check the answer against its sources with a Python review sheet.

A Tabstack Research call returns a Markdown report and the pages it cites. The Python example saves both, records the request’s progress events, and generates a CSV review sheet that connects each report sentence to its cited pages.

Use the timeline to inspect the request’s event order and elapsed time. Use the review sheet to check whether a source supports the claim attached to it. For example, the recorded answer describes a search result as page content, while the cited documentation specifies a content snippet. That distinction affects what your application can do with the result.

This tutorial uses the committed Ollama web-search example. You will run the CLI, inspect its output files, and follow one sentence through the source review. The implementation extends Build a cited live-web answer in Python with one Research call.

Run the traced request

You need Python 3.9 or later, uv, and a Tabstack API key. The example’s lockfile pins the tabstack SDK at 2.8.5. Run the following commands in Bash:

git clone https://github.com/Mozilla-Ocho/tabstack-cited-research-python.git
cd tabstack-cited-research-python
uv sync --frozen

# Read the key without echoing it or saving it to shell history.
read -rs TABSTACK_API_KEY && export TABSTACK_API_KEY

uv run cited-research \
  --query "What are the current ways to add web search to an Ollama-based application, and what output does each approach return?" \
  --mode fast --nocache --silence-timeout 120 \
  --output artifacts/my-trace

The CLI reads TABSTACK_API_KEY from the environment. The command has no API-key argument. --nocache bypasses the fetched-page cache, and --silence-timeout 120 stops the client after 120 seconds without an event.

The command writes your files to artifacts/my-trace. The committed example below uses artifacts/trace-run; its paths identify that saved run, not the directory you just selected.

The recorded run on 29 September 2026 produced this terminal output:

start Starting research
planning:start Planning research strategy
planning:end Planning complete: 6 queries
iteration:start iteration 1/1 Starting iteration 1 of 1
searching:start iteration 1 Searching with 6 queries
searching:end iteration 1 6 new urls Found 6 URLs
iteration:end iteration 1 Iteration 1 complete (fast mode)
writing:start Writing report
writing:end Report draft complete
complete report -> artifacts/trace-run/report.md
sources (3) -> artifacts/trace-run/sources.json
  1. Web search https://docs.ollama.com/capabilities/web-search
  2. Web search https://docs.ollama.com/capabilities/web-search.md
  3. Subagents and web search in Claude Code https://ollama.com/blog/web-search-subagents-claude-code
review sheet -> artifacts/trace-run/review-sheet.csv (unreviewed)

The first event arrived after 496 ms, and the request completed in 17.1 seconds. These timings describe the recorded request; they are not a general latency benchmark.

Open the output files

The example saves the answer, source list, and timeline separately so you can inspect each without reconstructing the stream:

artifacts/trace-run/
  question.txt
  command.txt                 the command as run, no secrets
  events.sanitized.jsonl      the lifecycle timeline
  report.md                   the answer
  sources.json                cited pages in returned order
  review-sheet.csv            one row per sentence, ready for review
  trace-diagram.md            the timeline as a diagram, plus its limits
  run-manifest.json           versions, UTC times, terminal state, caveats

Start with report.md, sources.json, and review-sheet.csv. The report contains the answer’s numbered citation markers. The source list preserves their positions. The review sheet joins sentences to those entries and leaves evidence and scoring fields for you to complete.

The sequence diagram shows where that review fits:

sequenceDiagram
    participant U as You
    participant C as cited-research CLI
    participant T as Tabstack Research
    participant V as Reviewer
    U->>C: Public question, fast mode
    C->>T: One POST /research (streamed)
    T-->>C: Progress events
    T-->>C: complete (report + cited pages) or error
    C-->>U: Timeline, report, cited pages, review sheet
    U->>V: Report and source record
    V-->>U: Claim-by-claim decisions and gaps

Preserve the citation mapping

sources.json numbers cited pages from 1 in the order the API returns them. Each record keeps the source’s id, url, title, claims, and source_queries, along with the CLI’s validation flags. The sample returned a title for every entry; it did not populate the optional relevance or reliability fields.

The report’s [n] markers refer to the source at position n. Sorting the list or removing an entry changes the targets of subsequent markers. The CLI preserves source order so each citation number continues to point to the correct page.

The first run returned three entries, but entries 1 and 2 point to the same Ollama documentation page, with and without .md. The CLI retains both entries and flags the repeated source.

Its duplicate detector compares URL patterns, including scheme, www., trailing slashes, and .md suffixes. Those comparisons flag possible matches; they do not inspect page contents. The sheet’s “[2] likely same page as [1]” label records that heuristic. In this example, both URLs resolve to the same documentation page.

The CLI also checks URLs before rendering them as links. It excludes loopback and private addresses, non-HTTP schemes, and URLs containing embedded credentials, and strips credentials from stored URLs.

Fast mode returns an empty claims array for each cited page. Use the report’s numbered markers to follow its statements to their source entries. Balanced mode populates the claim lists, as the Research guide documents. Those lists record the pipeline’s extracted claims; the review still checks them against the page.

If the response contains no cited pages, the CLI writes [] to sources.json and records review_needed_no_sources in the manifest.

Check a sentence against its source

The review sheet starts with one row per report sentence. It fills the citation mapping and reserves the evidence fields for the reviewer:

Fields Who fills them
claim_id, answer_text, citation_ids, source_url The CLI
passage, source_date_or_version, retrieved_at_utc, support, reason You, during source review

Helper columns include cited page IDs, automatic flags, and any claims supplied by the API. The CLI leaves support blank. Split a sentence into separate rows when its clauses make claims that need different evidence.

In the recorded answer, one sentence describes results containing title, url, and content “of a relevant web page” and cites entries [1][2]. Both entries lead to the same documentation. The relevant passage defines the field as:

content (string): relevant content snippet from the web page

The answer omits “snippet”. Correct that clause to say the result contains a relevant content snippet from the page. An application that needs the full page must fetch it rather than treat the search snippet as complete page content.

Record the passage, the source URL, when you inspected it, and why the wording needs correction. The citation review method defines the support values: 2 for supported at the stated scope, 1 for partial support or a missing qualifier, 0 for unsupported or contradicted after inspection, and U for uninspected or inaccessible evidence.

The original review also identified these specific issues:

Answer detail Source evidence Correction
An API key requires a free Ollama account, but the sentence has no inline citation “A free Ollama account is required.” Attach the supporting documentation to that claim.
“MCP (Multi-Context Processor)” The documentation describes a Python MCP server; MCP stands for Model Context Protocol. Replace the incorrect expansion with “Model Context Protocol”.
Built-in search applies to “specific” cloud models “It works with any model on Ollama’s cloud.” Do not turn recommended models into a restriction on supported models.

Keep missing-citation and duplicate-source flags separate from factual support scores. A repeated URL describes the source list; it is not another factual claim. A sentence without a marker needs an evidence check before you classify it as unsupported.

The question also asks what each approach returns. Check that required element separately for every approach; correcting one sentence does not establish whole-answer coverage. Write down those required elements before reading a new answer.

cited-research-review preserves completed review entries and edited rows. It will not regenerate over them unless you pass --force.

Read the event timeline

events.sanitized.jsonl records one event per line using an allowlist of fields. The CLI redacts URLs, email addresses, and credential-like strings from status messages, and omits payload fields such as planned queries.

This event comes from the second recorded run:

{"elapsed_ms": 3091, "event": "iteration:end", "is_last": true, "iteration": 1.0, "known_event": true, "message": "Iteration 1 complete (fast mode)", "received_at_utc": "2026-09-29T18:58:40.193Z", "seq": 7, "stop_reason": "max_iterations", "timestamp_raw": 1790708320152.0, "timestamp_type": "float"}

Use seq to preserve arrival order. Several events share a server timestamp, including start and planning:start, so timestamps alone cannot distinguish their sequence.

The client measures elapsed_ms from just before it sends the request. That value records the wait your application experienced, not the server’s internal execution time. The CLI also keeps timestamp_raw and its type. The sample carries a float in Unix epoch milliseconds; the current Research guide documents epoch milliseconds, and the API reference defines the field as a number.

The first recorded run produced this timeline:

seq Event Elapsed
1 start 496 ms
2 planning:start 496 ms
3 planning:end 1,565 ms
4 iteration:start 1,566 ms
5 searching:start 1,566 ms
6 searching:end 6,843 ms
7 iteration:end 6,843 ms
8 writing:start 6,843 ms
9 writing:end 17,083 ms
10 complete 17,084 ms

In this run, the client observed roughly 1.1 seconds between the planning events, 5.3 seconds between the searching events, and 10.2 seconds between the writing events. Use these events to update a progress indicator while the request runs.

The stream exposes request phases, not individual page-fetch timings or internal reasoning. The complete event supplies the final report and cited-page list; it does not expose the passages the service read.

Handle failures without duplicating requests

Except when TABSTACK_API_KEY is missing, the CLI records a terminal state in the manifest and returns a corresponding exit code:

Exit Meaning
0 complete arrived and every file was written
2 The stream sent an error event
3 The request was rejected before streaming began (401, 429, 5xx)
4 Connection or transport failure before the stream opened
5 TABSTACK_API_KEY is not set; no request was made
6 The stream closed before complete or error
7 No event arrived for --silence-timeout seconds
8 A second complete or error arrived after complete
9 The connection dropped or timed out after the stream opened
10 complete arrived without a report
11 Any other unexpected failure; the files are still written

Disable automatic retries

The CLI disables SDK retries and adds no retry loop. A connection failure can occur after the server accepts a request. Retrying automatically could create another research task and incur another charge. Inspect the failure before deciding whether to send a new request.

The first recorded run predates this change and records sdk_max_retries: 2. The second records 0. Check the manifest when comparing saved runs so you know which retry behavior applied.

Measure stream silence

The Research API has no overall server-side timeout. The CLI resets its silence timer on every event, including the wait for the first event. A request can continue as long as events arrive within the configured interval.

A silence timeout stops the client from waiting; it does not establish that the server cancelled the task. The CLI reports that distinction on exit.

The example includes offline fixtures for clean completion, absent sources, streamed errors, premature closure, duplicate terminal events, HTTP and transport failures, silence timeouts, and a missing API key. These tests exercise client handling without sending Research requests.

Use the trace in your application

Use a managed Research call when your application needs the finished answer and its source record without operating the research loop. Keep orchestration in your own application when you need to control individual searches, fetches, and prompts, or change direction between steps. A search API fits workflows that need links and snippets for the application’s own research process.

Your application sets the acceptance rules for the returned answer. Check decision-driving claims against their sources and choose how to handle missing evidence. The trace supplies the request record and citation mapping; the review supplies the support decisions.

Research runs on hosted infrastructure and sends content and instructions to contracted model providers. Use public, non-sensitive questions for this example. Read What “private by default” means for a managed web API and the Data Handling documentation for the processing path and settings, with the Privacy Notice as the canonical policy.

The trace does not record billing usage. Check your console for the request’s cost.

Reproduce and review the example

The original sample records Python 3.12.13, uv 0.11.28, and macOS arm64. Its reproducibility references are commit 7fe5f20 and the manifest’s 82b9be9, the same tree before a commit-message change. Keep those recorded versions and commit references when comparing the historical artifacts with a new run.

Run the traced Python example with one public question your application answers. Open report.md beside review-sheet.csv, follow one citation to its passage, and record your decision. Review the remaining material claims and required answer elements before using the answer downstream.

Use What makes a citation useful in an AI-generated answer? for the review rubric. The Research guide and API reference define the stream and response fields.

START FREE

Read the guide, then make the call.

Start with 10,000 free credits. No credit card required.

curl -fsSL https://tabstack.ai/install.sh | sh