> **For coding agents and LLMs:** This is one page from the Social Fetch docs (markdown export). For curated orientation and workflow guidance, start with [`/llms.txt`](https://www.socialfetch.dev/llms.txt); for agent onboarding and crawl rules, use [`/agents.txt`](https://www.socialfetch.dev/agents.txt); for the full endpoint list with links to pages like this one, use [`/llms-endpoints.txt`](https://www.socialfetch.dev/llms-endpoints.txt); for one platform's parameters and curls, use [`/llms-{platform}.txt`](https://www.socialfetch.dev/llms-tiktok.txt); use [`/llms.json`](https://www.socialfetch.dev/llms.json) when you need structured JSON for tool registration.

## This page

- **On-site (HTML):** [https://www.socialfetch.dev/docs/api/v1/web/extract/post](https://www.socialfetch.dev/docs/api/v1/web/extract/post)
- **Markdown (.mdx) URL:** [https://www.socialfetch.dev/docs/api/v1/web/extract/post.mdx](https://www.socialfetch.dev/docs/api/v1/web/extract/post.mdx)

## API base URL and authentication

- **API origin (from OpenAPI `servers`):** `https://api.socialfetch.dev`
- **Authentication:** send `x-api-key: sfk_...` on `/v1/**` routes unless the operation is explicitly anonymous (check OpenAPI `security`, the [API reference hub](https://www.socialfetch.dev/docs/api.mdx), [`/llms.txt`](https://www.socialfetch.dev/llms.txt), or [`/llms.json`](https://www.socialfetch.dev/llms.json) for each route).
- **OpenAPI JSON:** [https://www.socialfetch.dev/openapi.json](https://www.socialfetch.dev/openapi.json)

## Recommended docs entrypoints (this site)

- [Documentation overview](https://www.socialfetch.dev/docs.mdx) — top-level orientation (markdown).
- [Quickstart](https://www.socialfetch.dev/docs/quickstart.mdx) — authenticate with `x-api-key`, validate auth with `whoami`, and understand the JSON envelope.
- [SDK](https://www.socialfetch.dev/docs/sdk.mdx) — official TypeScript SDK guide, including `SocialFetchClient`, `Result`, and `unwrap()`.
- [SDK reference](https://www.socialfetch.dev/docs/sdk-reference.mdx) — exhaustive SDK method inventory and route mapping for agents, tooling, and power users.
- [Choose the right endpoint](https://www.socialfetch.dev/docs/choose-endpoint.mdx) — task-oriented route selection for smoke tests, profiles, list endpoints, and single-item lookups.
- [Capability matrix](https://www.socialfetch.dev/docs/capability-matrix.mdx) — fast comparison of identifiers, pagination, outcomes, media download, and SDK coverage.
- [Recipes](https://www.socialfetch.dev/docs/recipes.mdx) — copyable workflows (brand monitoring, transcripts, Ad Library, creator scoring, Reddit research) with credit callouts and SDK examples.
- [Integrations](https://www.socialfetch.dev/docs/integrations.mdx) — MCP for AI clients, n8n verified node, Apify Store Actors, Make custom app, SDK, and REST API connection paths.
- [MCP product page](https://www.socialfetch.dev/mcp) — hosted MCP overview, OAuth, Skills install.
- [MCP integration](https://www.socialfetch.dev/docs/integrations/mcp.mdx) — hosted `/mcp` server, OAuth, Cursor/VS Code/Claude install snippets, 165 endpoint tools, plus docs_search/docs_read for implementation help.
- [n8n integration](https://www.socialfetch.dev/docs/integrations/n8n.mdx) — install `n8n-nodes-socialfetch`, credentials, and workflow examples.
- [Apify integration](https://www.socialfetch.dev/docs/integrations/apify.mdx) — Store Actors under @social-fetch, PPE billing, dataset export, and quick start.
- [Make integration](https://www.socialfetch.dev/docs/integrations/make.mdx) — custom app modules for Make scenarios, API key credentials, and module catalog.
- [`/llms-endpoints.txt`](https://www.socialfetch.dev/llms-endpoints.txt) — every documented operation with a direct link to that route's agent-readable markdown page (prefer this over parsing OpenAPI).
- [`/llms-{platform}.txt`](https://www.socialfetch.dev/llms-tiktok.txt) — per-platform endpoint files generated from OpenAPI (parameters, credits, curls).
- [`/agents.txt`](https://www.socialfetch.dev/agents.txt) — agent crawl/onboarding file with capabilities, auth rules, and allowlist.
- [`/llms.json`](https://www.socialfetch.dev/llms.json) — structured machine-readable operation inventory with parameter names, pagination, outcomes, credits, and SDK mapping.
- [API reference hub](https://www.socialfetch.dev/docs/api.mdx) — human-friendly index of operations with links into generated pages.
- [Errors](https://www.socialfetch.dev/docs/errors.mdx) — shared error envelope and HTTP status guidance.
- [Credits](https://www.socialfetch.dev/docs/credits.mdx) — metering, `402`, and planning batch jobs.
- Outcome semantics such as `found`, `not_found`, and `private` are documented in [Errors](https://www.socialfetch.dev/docs/errors.mdx) and on operation pages when present in the OpenAPI contract.

## Markdown docs convention

- Every docs page has a markdown twin: append **`.mdx`** to the docs pathname (for example `/docs/quickstart` → `/docs/quickstart.mdx`).
- Agents that send `Accept: text/markdown` on `/docs/**` HTML URLs may receive markdown directly (same URL, `Vary: Accept`).

---
# Extract structured data from a web page (https://www.socialfetch.dev/docs/api/v1/web/extract/post)

## Summary

Extract structured fields from a web page using a CSS selector schema.

**Tags:** `Web`

## HTTP

- **Method:** POST
- **Path:** `/v1/web/extract`
- **Base URL:** `https://api.socialfetch.dev`

## Capability summary

- **SDK mapping:** `client.web.extract({ url: "https://example.com/products", schema: { name: "products", baseSelector: "div.product", fields: [{ name: "name", selector: "h2", type: "text" }] } })`
- **Pagination:** none
- **Business outcome field:** `data.lookupStatus` with values `found`, `restricted`

## Agent notes

- **Empty results:** `lookupStatus: restricted` means bot/access protection blocked the fetch; `extracted` is null.

## Credits

- **Base:** 2 credits per successful lookup.
- **Maximum on success (200):** 2 credits.
- **Normalization failure (502):** 2 credits charged.
- **Authoritative field:** `meta.creditsCharged`.

## Authentication

- **`x-api-key`**: API key (`sfk_...`)

## Request body

- **Required:** yes; content type `application/json`

**Shape:**

- **url** (required) — type `string`; minLength: 1; maxLength: 2083. Web page URL to fetch.
- **schema** (required) — type `object`. Crawl4AI JsonCssExtractionStrategy schema: baseSelector plus fields.
  - **name** (required) — type `string`; minLength: 1; maxLength: 200
  - **baseSelector** (required) — type `string`; minLength: 1; maxLength: 500
  - **fields** (required) — type `array`
    - _items:_
      - **name** (required) — type `string`; minLength: 1; maxLength: 100
      - **selector** (required) — type `string`; minLength: 1; maxLength: 500
      - **type** (optional) — type `string`; enum: text, attribute, html, regex; default: `text`
      - **attribute** (optional) — type `string`; minLength: 1; maxLength: 100
- **scanFullPage** (optional) — type `boolean`. When true, scroll the page to load dynamically appended content.
- **waitFor** (optional) — type `string`; maxLength: 200. Wait for a CSS selector before extraction. Must be prefixed with "css:" (e.g. css:main). JavaScript wait conditions are not supported.

**Example:**

```json
{
  "url": "string",
  "schema": {
    "name": "string",
    "baseSelector": "string",
    "fields": [
      {
        "name": "string",
        "selector": "string",
        "type": "text",
        "attribute": "string"
      }
    ]
  },
  "scanFullPage": false,
  "waitFor": "string"
}
```

## Responses (status codes)

- **200**: Structured CSS extraction result.
- **400**: Invalid body or disallowed URL
- **401**: Missing or invalid API key
- **402**: Insufficient credits
- **500**: Unexpected or billing error
- **502**: Extraction could not be completed.
- **503**: Service temporarily unavailable; safe to retry with backoff.

## Response body (200)

Structured CSS extraction result.

### Field outline

- **data** (required) — type `object`. Endpoint-specific response payload.
  - **lookupStatus** (required) — type `string`; enum: found, restricted. Whether page content could be extracted. Restricted means bot protection or similar access controls blocked automated fetching.
  - **url** (required) — type `string`. URL that was fetched.
  - **status** (required) — type `integer`; nullable. HTTP status code reported for the page fetch when available; null when restricted.
  - **extracted** (required) — type `array`; nullable. Structured rows extracted via CSS schema when lookupStatus is found; null when restricted.
    - _items:_
      - type `object`
  - **metadata** (optional) — type `object`. Page metadata such as title when available.
- **meta** (required) — type `object`. Metadata describing the request and billing outcome.
  - **requestId** (required) — type `string`; minLength: 1. Unique request identifier for tracing this API call.
  - **creditsCharged** (required) — type `integer`; minimum: 0. Credits charged for this request.
  - **version** (required) — type `string`; enum: v1. Public API version that served the response.
  - **cached** (optional) — type `boolean`. True when served from shared response cache. Credits still apply (full endpoint price); Age header may be present.

### Example JSON (OpenAPI example)

```json
{
  "data": {
    "lookupStatus": "found",
    "url": "https://example.com/products",
    "status": 200,
    "extracted": [
      {
        "name": "Widget",
        "price": "$9.99"
      }
    ],
    "metadata": {
      "title": "Products"
    }
  },
  "meta": {
    "requestId": "req_web_extract_example",
    "creditsCharged": 2,
    "version": "v1"
  }
}
```

### Machine-readable error codes

When an error JSON body is returned, it may include one of these `error.code` values (derived from the OpenAPI schemas for this operation; additional codes may exist at runtime):

- `bad_request`

## Error handling & retries

Interpret HTTP status codes using the descriptions below. Do not assume a JSON body unless the OpenAPI schema defines one for that status.

- **400**: Invalid body or disallowed URL **Retry:** Fix the request; retrying the same invalid payload will not help.
- **401**: Missing or invalid API key **Retry:** Fix the API key first; retrying without changes will not help.
- **402**: Insufficient credits **Retry:** Do not retry without resolving billing/credits (retrying the same request will not help).
- **500**: Unexpected or billing error
- **502**: Extraction could not be completed. **Retry:** May be transient; a few retries with backoff are reasonable.
- **503**: Service temporarily unavailable; safe to retry with backoff. **Retry:** Usually safe to retry with exponential backoff and jitter.

### Suggested client defaults

- Send the API key using the `x-api-key` header on every request.
- On `503` (and sometimes `502`), retry with backoff; cap retries and surface a clear error to the user.
- On `402`, surface an actionable billing message rather than blind retries.

## Examples

### TypeScript SDK

```typescript
import { SocialFetchClient } from "@socialfetch/sdk";

const client = new SocialFetchClient({
  apiKey: process.env.SOCIALFETCH_API_KEY!,
});

const result = await client.web.extract({
    url: "https://example.com/products",
    schema: {
      name: "products",
      baseSelector: "div.product",
      fields: [
        {
          name: "name",
          selector: "h2",
          type: "text"
        },
        {
          name: "price",
          selector: ".price",
          type: "text"
        }
      ]
    }
  });

if (!result.ok) {
  console.error(result.error);
} else {
  console.log(result.value.data);
}
```

### Node.js

```javascript
const response = await fetch(
  "https://api.socialfetch.dev/v1/web/extract",
  {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      "x-api-key": "YOUR_API_KEY"
    },
    body: JSON.stringify({
      url: "https://example.com/products",
      schema: {
        name: "products",
        baseSelector: "div.product",
        fields: [
          {
            name: "name",
            selector: "h2",
            type: "text"
          },
          {
            name: "price",
            selector: ".price",
            type: "text"
          }
        ]
      }
    })
  }
);

const data = await response.json();
console.log(data);
```

### cURL

```bash
curl "https://api.socialfetch.dev/v1/web/extract" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/products","schema":{"name":"products","baseSelector":"div.product","fields":[{"name":"name","selector":"h2","type":"text"},{"name":"price","selector":".price","type":"text"}]}}' \
  -X POST
```

### Python

```python
import requests

response = requests.post(
    "https://api.socialfetch.dev/v1/web/extract",
    headers={"x-api-key": "YOUR_API_KEY", "Content-Type": "application/json"},
    json={"url":"https://example.com/products","schema":{"name":"products","baseSelector":"div.product","fields":[{"name":"name","selector":"h2","type":"text"},{"name":"price","selector":".price","type":"text"}]}},
)
data = response.json()
print(data)
```