Extract
Turn a URL into clean text or matching JSON.
Pass a page and the shape your application needs. Tabstack returns usable output without a parser, browser fleet, or second model call in your stack.
Clean Markdown · Schema-matched JSON · Cache, effort and geo controls
TEXT
MARKDOWN
LINKS
const res = await client.extract.json({
url: 'https://example.com/pricing',
json_schema: {
type: 'object',
properties: {
plan: { type: 'string' },
price: { type: 'number' },
},
},
})res = client.extract.json(
url='https://example.com/pricing',
json_schema={
'type': 'object',
'properties': {
'plan': {'type': 'string'},
'price': {'type': 'number'},
},
},
)tabstack extract json https://example.com/pricing \
--schema '{"type":"object","properties":{"plan":{"type":"string"},"price":{"type":"number"}}}'The Extract API
What does the Tabstack Extract API return?
Tabstack Extract takes a URL and returns clean Markdown or JSON matching the schema you provide. Fetching, rendering, content extraction, and schema enforcement happen before the response comes back.
tabstack
fetch, render, extract
enforce the schemaChoose the output
Read the page or return the object.
/extract/markdown
Use it when your model, vector store, or application needs readable page content with headings and links intact.
/extract/json
Use it when the next step expects specific fields and types. Provide the schema with the request and receive matching output.
Inside the call
Four things happen before the response comes back.
Fetch the page
Retrieve the URL you passed, with the cache and geo controls you set.
Render what needs it
Run JavaScript when the content needs it.
Extract the content
Separate the page from its chrome and keep the structure.
Enforce the schema
Return the fields and types your request asked for.
const res = await client.extract.json({
url: 'https://example.com/pricing',
json_schema: {
type: 'object',
properties: {
plan: { type: 'string' },
price: { type: 'number' },
seats: { type: 'number' },
},
},
})
// every key you ask for is present; validate the valuesres = client.extract.json(
url='https://example.com/pricing',
json_schema={
'type': 'object',
'properties': {
'plan': {'type': 'string'},
'price': {'type': 'number'},
'seats': {'type': 'number'},
},
},
)
# every key you ask for is present; validate the valuestabstack extract json 'https://example.com/pricing' \
--schema '{"type":"object","properties":{"plan":{"type":"string"},"price":{"type":"number"},"seats":{"type":"number"}}}'Schemas
You define the schema. Tabstack returns the data.
Missing fields: Every key you ask for is present. A field the page does not state can come back null, empty, a placeholder number, or a guessed value. Validate values, not just keys.
In your pipeline
Keep extraction out of your application code.
Page layouts change. Rendering requirements vary. A field that looks simple can require navigation, retries, or model reasoning. Tabstack keeps those concerns behind the API while your application keeps a stable output contract. Use Extract for known URLs and repeated jobs.
your code
one call
one shapebehind the api
fetch, render
retries, reasoningChoosing the call
Extract returns what the page says. Generate works out what it means.
- Need the fields that are on the page? This endpoint returns them in your schema.
- Need interpretation, classification, or comparison? Generate adds instructions and returns that result in your schema.
- Extract the plan names and prices, then ask Generate which audience each plan appears to target, and why.
Controls
Match the call to the page and the job.
Three parameters change how the page is retrieved.
| Control | Use it when |
|---|---|
nocache | The workflow cannot reuse a cached result |
effort | You want to choose the retrieval depth |
geo_target | The page varies by country |
Use cases
Use Extract when a known page feeds another step.
Read a pasted link
Turn a URL someone dropped in as clean, readable Markdown.
Current docs for an assistant
Give a coding assistant documentation newer than its training data.
Keep a source set current
Refresh a defined set of pages that a retrieval index depends on.
Enrich an incoming record
Fill out a company or contact record from its public pages.
Monitor a public page
Track product, price, job, policy, or filing pages over time.
Typed records for a data product
Turn pages into rows your own product can query and ship.
Trust
How Extract handles your data.
Never trained on. Private by default. The Trust page explains what is processed and the account controls.
tabstack
processes the page
returns the outputnever trained on
private by default