Source: http://www.poma-ai.com/docs/grill/concepts/ingestion

# Grill Ingestion

Grill's ingestion entry point is **`POST /grill/ingest`**. It runs the same PrimeCut pipeline — parse, chunk, embed — but **persists the result inside your project's namespace** instead of returning a `.poma` archive. Once a job finishes the document is immediately searchable through `/grill/search`; you never download or re-upload anything.

## Why a separate endpoint

`POST /grill/ingest` and `POST /primeCut/ingest` look almost identical on the wire. They differ in what happens **after** the chunks are produced:

| Step | `/primeCut/ingest` | `/grill/ingest` |
| --- | --- | --- |
| Parse + chunk | ✅ same pipeline | ✅ same pipeline |
| Embed | Optional, depends on plan | ✅ always |
| Persist to project namespace (vectors + storage) | ❌ | ✅ |
| Make doc available to `/grill/search` | ❌ | ✅ |
| `.poma` archive download via `/jobs/{job_id}/download` | ✅ | ❌ |

If you want chunks to take home, use PrimeCut. If you want the document **searchable** through `/grill/search`, use Grill.

> One project = one product. A project created with `product:"primecut"` cannot call `/grill/ingest`, and vice versa. See [Create a Grill project](/grill/getting-started/projects).

## The job lifecycle

Ingestion is asynchronous. The `job_id` you get back moves through the standard POMA lifecycle:

```text
pending  ──▶ processing  ──▶ done       ◀── searchable from this point
                          └─▶ failed
```

At `done`, three things are true atomically: the document appears in `GET /grill/docs`, `GET /grill/docs/{docId}` returns its `DocInfo`, and `/grill/search` can retrieve from it. The SDK's `ingest()` **waits for this for you**; with raw HTTP you poll `GET /jobs/{job_id}/status` or stream `GET /status/v1/jobs/{job_id}`. There's no `.poma` to download — Grill keeps the artifacts server-side.

`docId` is derived from the filename (sanitised) plus a project salt; read the canonical value from `DocInfo.doc_id` and use it for `doc_filter` and `/grill/docs/{docId}`.

## Ingest with the Python SDK

The recommended path. The `Grill` client handles the octet-stream framing, polling, and status streaming for you, and reads the project API key from `POMA_GRILL_API_KEY`.

```python
# pip install poma
from poma import Grill

g = Grill()
result = g.ingest("manual.pdf")        # submit + wait, returns when done
print(result.job_id, result.status, result.usage)
```

**Batch** — submit everything first, collect later so the waits overlap:

```python
g = Grill()
job_ids = [g.submit(p) for p in ["a.pdf", "b.pdf", "c.pdf"]]
results = [g.collect(jid) for jid in job_ids]
```

**Async** — run the waits concurrently with `AsyncGrill`:

```python
import asyncio
from poma import AsyncGrill

async def ingest_all(paths: list[str]) -> None:
    async with AsyncGrill() as g:
        results = await asyncio.gather(*(g.ingest(p) for p in paths))
        for r in results:
            print(r.job_id, r.status)

asyncio.run(ingest_all(["a.pdf", "b.pdf", "c.pdf"]))
```

**From a URL** instead of a local file — pass `remote_url` and the server fetches it:

```python
g.ingest(remote_url="https://example.com/report.pdf")
```

Full signatures: [`Grill` reference](/sdk/reference/grill), [`AsyncGrill` reference](/sdk/reference/async-grill).

## Attributes & metadata for filtering {#attaching-metadata-for-filtering}

Tag a document at ingest so queries can filter on it later. For all new work, use **typed attributes**: one map that holds any type.

### Check what already exists first

Before your code — or your agent — sends an attribute name, **list the ones the project already has** with `GET /grill/attributes` (MCP: [`grill_attributes`](/mcp/grill-mcp#grill-attributes)). A name and its type are permanent once used, and each project has a budget of 64 names, so:

1. **Reuse** an existing name and type where one fits. If `year` (`int`) exists, don't add `doc_year`, and don't send `"2025"` as a string to it.
2. **Declare a new name only when nothing fits**, and pick it deliberately — it can never be renamed or deleted.
3. **Treat a `503` as "retry later", never as "no attributes exist".** Acting on an empty list you didn't actually get is how a project ends up with `region`, `region_code` and `doc_region` side by side.

```bash
curl -sS "$GRILL/attributes" -H "authorization: Bearer $GRILL_KEY"
# {"attributes":[{"name":"region","type":"string"},{"name":"year","type":"int"}],"max_names":64,"note":"…"}
```

A typical agent flow:

```text
1. GET /grill/attributes              → region:string, year:int      (503 → back off and retry)
2. The new document has a region, a year, and a reviewer's summary.
   region, year  → reuse as-is (same names, same types)
   review_notes  → nothing fits; new name, declare it as encrypted_text
3. POST /grill/ingest
     X-Attributes:       {"region":"emea","year":2025,"review_notes":"…"}
     X-Attribute-Schema: {"review_notes":{"type":"encrypted_text"}}
```

This matters most for agents: an LLM that invents a fresh name per document burns the name budget in a few dozen ingests and splits your filters across synonyms.

### Typed attributes

```python
grill.ingest(
    "report.pdf",
    attributes={
        "region": "emea",            # string
        "revision": 12,              # int
        "reviewed": True,            # bool
        "published_at": "2026-09-10T00:00:00Z",
        "case_notes": "patient presented with chest pain",
    },
    attribute_schema={
        "published_at": {"type": "datetime"},
        "case_notes": {"type": "encrypted_text"},
    },
)
```

`string`, `int`, `float`, `bool` and their arrays are inferred from the value.
`datetime` and `encrypted_text` must be declared — pass `attribute_schema=` in
the SDK (the `X-Attribute-Schema` header over HTTP) — because neither is
distinguishable from a plain string by looking at it; an undeclared new name is
stored as a plain `string`. Over HTTP both maps travel as headers
(`X-Attributes`, `X-Attribute-Schema`), so keep them under 2 KB each. Names must
match `^[a-z0-9_]{1,64}$`.

**`encrypted_text` is the one nobody else offers.** It gives you full-text
token search over content the vector database cannot read: the tokens are
hashed with your tenant key before they leave us, and the database matches
hashes it cannot reverse. Use it for case notes, contract clauses, patient
summaries — anything you want searchable but are not allowed to expose.

Four things to know before you write the first one:

- **A type is permanent.** The first document that uses a name fixes its type
  for the whole project, and it can never be changed. Send `12`, not `"12"`.
- **A name is permanent too**, and each project has a budget of 64. Clearing a
  value later does not free its name, so a typo costs a slot forever — which is
  why you [check the existing names first](#check-what-already-exists-first).
- **`encrypted_text` matches exact tokens.** No stemming, no partial words, no
  wildcards — a hash has no prefix. `"pains"` will not find `"pain"`.
- **Filters are AND.** A filter naming an attribute your project has never used
  matches nothing, rather than being quietly ignored.

The HTTP API also accepts **`X-Unencrypted-Strings`** for plaintext, wildcard-filterable metadata (see the [API reference](/grill/reference/api)) — never put secrets or PII there.

### Migrating from labels / meta_int {#migrating-from-labels-meta-int}

Earlier versions attached metadata through three fixed fields: `labels`, `meta_int_1` and `meta_int_2`. They are now typed attributes with reserved names. The old fields are still accepted and mapped onto the new ones; their removal will be announced in advance. New code should send only `attributes` / `X-Attributes` and filter only with `attribute_filters`. The Python SDK still sends the old arguments as the old headers, and from version 0.7.0 it emits a `DeprecationWarning` when you use them.

| Old | New |
| --- | --- |
| SDK `labels=["q3 report", "emea"]` (warns from SDK 0.7.0) · header `X-Labels` | SDK `attributes={"labels": ["q3 report", "emea"]}` · header `X-Attributes: {"labels": [...]}` |
| SDK `meta_int_1=n` (warns from SDK 0.7.0) · header `X-Meta-Int-1` | `attributes={"unencrypted_int_1": n}` · `X-Attributes: {"unencrypted_int_1": n}` |
| SDK `meta_int_2=n` (warns from SDK 0.7.0) · header `X-Meta-Int-2` | `attributes={"unencrypted_int_2": n}` · `X-Attributes: {"unencrypted_int_2": n}` |
| SDK `labels_any=[…]` · search field `meta_tags_any` | `attribute_filters=[{"attr": "labels", "op": "has_any_tokens", "value": "q3 emea"}]` |
| SDK `labels_all=[…]` · search field `meta_tags_all` | `attribute_filters=[{"attr": "labels", "op": "has_all_tokens", "value": "q3 emea"}]` |
| `meta_int_1_gte` / `meta_int_1_lte` | `{"attr": "unencrypted_int_1", "op": "gte", "value": n}` / `"op": "lte"` |
| `meta_int_2_gte` / `meta_int_2_lte` | `{"attr": "unencrypted_int_2", "op": "gte", "value": n}` / `"op": "lte"` |

What changes, and what to watch for:

- **Types.** `labels` is `[]encrypted_text`: values stay HMAC'd with your tenant key, as before. `unencrypted_int_1` and `unencrypted_int_2` are plaintext `int`s, range-filterable as before. These types come with the reserved names, so — unlike other `encrypted_text` attributes — you do not declare them in `attribute_schema=` / `X-Attribute-Schema`; an undeclared `labels` is never inferred as a plain `[]string`.
- **Size.** Labels now travel inside `X-Attributes`, which is capped at 2048 characters per request together with every other attribute you send. A very large label set that fit the old `X-Labels` limits may need trimming.
- **Label matching is token-based.** The filter `value` is a string, matched against the words of each label rather than the whole label: a label `"q3 report"` stored as a typed attribute matches a filter on `"q3"`. Exact equality on a whole label is no longer available.
- **Documents tagged only through the old fields still match, but on whole tags.** For a document tagged only through the deprecated `labels=` / `X-Labels`, a `labels` filter matches when the whole filter value equals one of its tags, or when the value's whitespace-separated words equal whole tags — every word for `has_all_tokens`, any word for `has_any_tokens` (up to 32 words; a longer value is matched only as one whole tag). A single word from inside a multi-word legacy tag does not match: a document tagged `"q3 report"` only through `labels=` / `X-Labels` is found by `"q3 report"`, but not by `"q3"`.
- **The three names are reserved.** Using `labels`, `unencrypted_int_1` or `unencrypted_int_2` as an attribute name always means the semantics above, and each counts against the project's budget of 64 names once used.
- **Punctuation-only labels** (for example `"--"`) produce no tokens, so they cannot be filtered on as attributes.

See [Retrieval](/grill/concepts/retrieval) for how to apply these filters at search time.

## Direct API usage

If you're not on Python, `POST /grill/ingest` takes **raw file bytes** as `application/octet-stream`, with the filename in `Content-Disposition`:

```bash
curl -sS -X POST "$GRILL/ingest" \
  -H "authorization: Bearer $GRILL_KEY" \
  -H "content-type: application/octet-stream" \
  -H 'content-disposition: attachment; filename="manual.pdf"' \
  --data-binary @manual.pdf
```

- Multipart (`multipart/form-data`) is **not** accepted → `403`. Use octet-stream.
- Fetch from a URL instead of uploading bytes by setting `X-Remote-URL` (then `Content-Disposition` and the body are optional).
- Metadata rides as headers: `X-Attributes` / `X-Attribute-Schema` (typed attributes — [list existing names](#check-what-already-exists-first) with `GET /grill/attributes` first), `X-Unencrypted-Strings`, plus `X-Base-URL` and `X-Completion` (completion webhook).

Full request/response shapes and every header are in the [API reference](/grill/reference/api).

## Parallel ingests

`POST /grill/ingest` takes **one file per request** — there is no multi-file body (multipart is rejected). To ingest many files, submit them as parallel single-file requests. Per-account queue limits apply: up to **20 jobs processing concurrently** and **1,000 jobs queued** (free/demo accounts: 1 concurrent, 20 queued). Past the limit the API returns `429 too_many_jobs` with a `Retry-After` header — wait and resubmit; nothing is lost.

## Bulk ingest from a bucket

Have a whole corpus in S3 or GCS? `POST /grill/ingestBucket` ingests the bucket (or a prefix) as one server-side task — no per-file uploads, no request-size limit, up to 500 MB per file and 500 GB per task, one active task per account. See the [API reference](/grill/reference/api#post-grill-ingestbucket) for the request shape, task lifecycle, and all limits.

## Supported file types

Grill inherits the full PrimeCut format set:

- **Documents:** `pdf`, `doc`, `docx`, `dotx`, `rtf`, `txt`, `md`, `html`, `htm`, `xml`
- **Data & structured text:** `json`, `yaml`, `toml`, `ini`, `env`, `cir`
- **Presentations:** `ppt`, `pptx`, `pptm`, `pps`, `ppsx`, `pot`, `potx`, `key`
- **Spreadsheets:** `xls`, `xlsx`, `xlsm`, `xlsb`, `xltx`, `csv`, `tsv`, `numbers`, `ods`, `odc`
- **Images:** `png`, `jpg`, `jpeg`, `gif`, `bmp`, `tif`, `tiff`, `svg`, `webp`, `ico`, `heic`, `heif`, `psd`
- **Audio:** `mp3`, `wav`, `m4a`, `aac`, `ogg`, `flac`, `opus`
- **Video:** `mp4`, `mov`, `webm`, `mkv`, `avi`, `m4v`, `mpeg`, `mpg`
- **Other:** `epub`, `mobi`, `djvu`, `dwg`, `dxf`, `dwf`, `dwfx`, `vsd`, `vsdx`, `ai`, `eps`, `ps`, `prn`, `xps`, `oxps`, `pub`, `mdi`, `pages`, `odp`, `odf`, `odt`

**Audio & video** are handled by a native media front-end — video is keyframe-sampled and transcribed (speech-to-text) with per-keyframe vision, audio runs the same understanding pass without frames — so both become searchable text like any other document.

**Ingesting from a URL:** set `X-Remote-URL` to have Grill fetch the bytes instead of uploading them (the `Content-Disposition` and body become optional). Public video-platform URLs (e.g. YouTube) are supported as a source too.

## Re-ingesting the same file

Re-uploading a file with the **same effective `docId`** replaces the existing document — old vectors and storage are discarded. There's no append mode; an updated PDF fully supersedes the previous version. For version history, ingest each version under a distinct filename so the `docId` differs. To remove a doc cleanly first, use `DELETE /grill/docs/{docId}` ([Document management](/grill/concepts/document-management)).

## Errors you will see

| Status | When | What to do |
| --- | --- | --- |
| `400` | No `X-Remote-URL` and a missing/invalid `Content-Disposition`, unsupported MIME, or empty body | Fix the headers; check the file is non-empty (or supply `X-Remote-URL`). |
| `401` | Missing or invalid Bearer token | Use a project API key — see [Authentication](/grill/getting-started/authentication). |
| `403` | Caller's project is `primecut`, or the request is multipart | Create a Grill project; switch to octet-stream. |
| `413` | Upload body exceeds the **50 MB** request limit | Host the file and pass `remote_url` / `X-Remote-URL` instead of uploading the bytes. |
| `429` | Too many jobs queued/processing for the account (`too_many_jobs`) | Honor the `Retry-After` header and resubmit — or use [bucket ingest](/grill/reference/api#post-grill-ingestbucket) for large corpora. |
| `500` | Server-side parse failure | Retry once; if it persists, contact support with the `job_id`. |

## Next

- [Retrieval](/grill/concepts/retrieval) — once the doc is in, how search behaves.
- [Document management](/grill/concepts/document-management) — list, inspect, delete.
- [API reference](/grill/reference/api) — full `/grill/ingest` request/response.