Extract schema library
Start with the object your workflow needs.
Browse ready-made JSON Schemas for common public-web data jobs. Copy one into /extract/json, adapt the fields, and receive matching JSON from the page you provide.
What a schema is
A schema defines the response.
The schema describes the field names, types, nesting, and descriptions your application expects. Tabstack uses it to shape the Extract response so the next step receives a known object instead of page content it has to interpret.
TEXT
MARKDOWN
LINKS
const res = await client.extract.json({
url: 'https://example.com/pricing',
json_schema: {
type: 'object',
properties: {
plan: { type: 'string' },
price: { type: 'number' },
},
},
})res = client.extract.json(
url='https://example.com/pricing',
json_schema={
'type': 'object',
'properties': {
'plan': {'type': 'string'},
'price': {'type': 'number'},
},
},
)tabstack extract json https://example.com/pricing \
--schema '{"type":"object","properties":{"plan":{"type":"string"},"price":{"type":"number"}}}'Categories
Ten categories.
B2B Intelligence
Company profiles, competitors, funding, pricing pages, changelogs and tech stacks.
Browse B2B intelligenceE-commerce
Product listings, reviews, category results, marketplace sellers and grocery items.
Browse e-commerceJobs and Hiring
Job postings, salary data, hiring velocity, layoffs and leadership changes.
Browse jobs and hiringReal Estate
Residential and commercial listings, rentals, sold history, foreclosures and short-term stays.
Browse real estateFinance
Public filings, crypto assets, VC portfolios, angel deals and alternative assets.
Browse financeGovernment and Public Records
Business registrations, court filings, contract awards, permits and property tax records.
Browse government and public recordsHealthcare
Clinical trials, provider directories, drug formularies and device recalls.
Browse healthcareDeveloper Ecosystem
GitHub repositories and package registry entries for OSS and competitive research.
Browse developer ecosystemExtract or Generate
Use the one that matches the page.
- Does the page state the fields? Extract returns them in your schema.
- Does the output need classification, comparison or inference? Generate adds instructions and returns that result in your schema.