Source: http://www.poma-ai.com/docs/mcp/grill-mcp

# `poma-grill-mcp` — Grill MCP Server

`poma-grill-mcp` is POMA AI's Model Context Protocol server for the **Grill** context engine. It exposes nine tools covering the full Grill loop: ingest (sync, async, batch, resume), search, job status, document listing, project listing, and a self-describing `grill_explain`. Two implementations ship in the same repo — a stable Go binary and a Node/TypeScript package — both speaking the same tool surface against the same backend. POMA also hosts the server, so you can use Grill from any MCP client **without any local install**.

- **Repo:** [github.com/poma-ai/poma-grill-mcp](https://github.com/poma-ai/poma-grill-mcp)
- **Hosted endpoint:** `https://mcp.poma-ai.com` (OAuth 2.0 — no key in your config)
- **Implementations:** Go (stable, default; `go install` + Docker + build-from-source), Node/TypeScript (`npx`, `npm install -g`, build-from-source)
- **Transports:** stdio, streamable HTTP, Docker (Go only)
- **License:** MPL-2.0

> **This server does not expose PrimeCut tools.** For `primecut_*` tools (raw chunk JSON, `.poma` archive bytes), use [`poma-mcp`](/mcp/poma-mcp). The two servers can run side-by-side in the same client config.

## When to use this server

Pick `poma-grill-mcp` when:

- You want **prompt-ready RAG context** from Grill without writing a single line of glue code in your agent.
- You're ingesting **large files** and want the server to read them from disk via `file_path` instead of inflating tool calls with base64.
- You want **zero-install** access via POMA's hosted endpoint at `https://mcp.poma-ai.com`.
- You're building a backend service that uploads via HTTP — the Go binary has a dedicated `POST /ingest-upload` endpoint that accepts raw bytes or multipart with the same auth as MCP.

## 1. Get an API key

Sign up at [console.poma-ai.com](https://console.poma-ai.com), create a project with `product: "grill"` ([Create a project](/grill/getting-started/projects)), and capture the project API key returned at creation (prefix `poma_prod_gr_…`) — it scopes every call to that project implicitly.

An **account-level key** (prefix `poma_acc_…`) also works: pass `project_id` on each tool call, or set the `POMA_PROJECT_ID` env var, to select which project to operate on. Use `grill_projects` to list the projects your key can see.

Using the **hosted endpoint** (Option B below)? You don't need a key in your config at all — auth is OAuth 2.0 via your Console login.

## 2. Install (skip for hosted)

If you're going to use the **hosted** endpoint, you don't need to install anything — jump to step 3, Option B.

### Go binary

> **Homebrew / `go install` are not available yet** — the tap has no published formula and the module layout isn't proxy-installable. Build from source, or use the Node package / hosted endpoint below. (`brew install poma-ai/poma/poma` installs the POMA **CLI**, not this server.)

```bash
git clone https://github.com/poma-ai/poma-grill-mcp
cd poma-grill-mcp/go
go build -o bin/poma-grill-mcp .
realpath bin/poma-grill-mcp
```

### Node / TypeScript

Requires Node 20+. Published as `@poma-ai/poma-grill-mcp` on npm.

::: code-group

```bash [npx (no install)]
# Most MCP clients can invoke npx directly — no separate install needed
npx -y @poma-ai/poma-grill-mcp -input -
```

```bash [Global npm]
npm install -g @poma-ai/poma-grill-mcp
poma-grill-mcp -input -
```

```bash [From source]
git clone https://github.com/poma-ai/poma-grill-mcp
cd poma-grill-mcp/node
npm install
npm run build
node dist/index.js -input -
```

:::

## 3. Add it to your agent

You have three options. **Option C (Agent Plugin)** is the fastest — one command wires the hosted server **and** installs two ready-made skills. Use Option B for a plain MCP config without skills, Option A when you need a local binary (e.g. `file_path` ingest from your machine).

| | What you get | Pick it when |
| --- | --- | --- |
| **C — Agent Plugin** | Hosted endpoint + `grill-ingest`/`grill-search` skills, auto-updates | You use Claude Code (or any Agent-Plugins-capable tool) and want the batteries-included setup |
| **B — hosted MCP** | Hosted endpoint, plain MCP config | Your client speaks `type: "http"` but not Agent Plugins, or you want zero extras |
| **A — local binary** | stdio server on your machine | You need `file_path`/`/ingest-upload` for large local files, or an air-gapped/self-hosted setup |

### Option A — local binary over stdio {#option-a-local-binary-over-stdio}

#### Option A1: Go binary

::: code-group

```json [Claude Code · ~/.claude/claude.json]
{
  "mcpServers": {
    "poma-grill-mcp": {
      "command": "/full/path/to/poma-grill-mcp",
      "args": ["-input", "-"],
      "env": { "POMA_API_KEY": "your-api-key" }
    }
  }
}
```

```json [Claude Desktop · claude_desktop_config.json]
{
  "mcpServers": {
    "poma-grill-mcp": {
      "command": "/full/path/to/poma-grill-mcp",
      "args": ["-input", "-"],
      "env": { "POMA_API_KEY": "your-api-key" }
    }
  }
}
```

```json [Cursor · ~/.cursor/mcp.json]
{
  "mcpServers": {
    "poma-grill-mcp": {
      "command": "/full/path/to/poma-grill-mcp",
      "args": ["-input", "-"],
      "env": { "POMA_API_KEY": "your-api-key" }
    }
  }
}
```

:::

#### Option A2: Node via `npx` (no install)

```json
{
  "mcpServers": {
    "poma-grill-mcp": {
      "command": "npx",
      "args": ["-y", "@poma-ai/poma-grill-mcp", "-input", "-"],
      "env": { "POMA_API_KEY": "your-api-key" }
    }
  }
}
```

### Option C — Agent Plugin (one command, adds skills) {#option-c-agent-plugin}

POMA hosts the server at `https://mcp.poma-ai.com` (any path on that host works the same; older configs with `/grill` keep working). Auth is **OAuth 2.0**: on first connect your client gets a `401` challenge and walks you through a browser login at [console.poma-ai.com](https://console.poma-ai.com) — no API key in your config. MCP clients (Claude Code, Claude Desktop, Cursor, …) handle the whole flow automatically.

::: code-group

```json [Claude Code / Desktop]
{
  "mcpServers": {
    "poma-grill-mcp": {
      "type": "http",
      "url": "https://mcp.poma-ai.com"
    }
  }
}
```

```json [Cursor]
{
  "mcpServers": {
    "poma-grill-mcp": {
      "url": "https://mcp.poma-ai.com"
    }
  }
}
```

:::

Prefer a static key over the browser flow (CI, headless agents)? Pass it as a header instead — `"headers": { "x-api-key": "your-api-key" }` — and the server skips OAuth for that request.

> The hosted server runs on POMA's infrastructure, so it **cannot read paths on your laptop**. Use `file_base64` when calling it from a hosted client, or use a local binary (Option A) when you need `file_path`.

After saving, restart the client. Sample prompt:

> "Ingest `~/docs/contract.pdf` into POMA Grill using **file_path**, then search it for 'termination clause'."

## Tools

`poma-grill-mcp` exposes ten tools covering the full Grill loop. Both Go and Node implementations expose the same surface.

| Tool | What it does |
| --- | --- |
| `grill_ingest` | Starts ingest; returns `job_id` after upload. **Does not wait** for indexing to finish. |
| `grill_ingest_sync` | Ingest + wait for terminal status. Returns `job_id` and the full `events` stream. |
| `grill_ingest_resume` | Reconnect to an in-progress job's status stream and wait until terminal — without re-uploading. |
| `grill_ingest_batch` | Upload up to 50 files with controlled concurrency. Returns `job_ids` as uploads complete. |
| `grill_jobs_status` | Snapshot the status of up to 50 jobs in one call. No streaming. |
| `grill_search` | Hybrid search returning prompt-ready context. Use `doc_filter` to scope to one doc, `attribute_filters` to filter on typed attributes. |
| `grill_attributes` | List the project's typed attributes (name + type) and the name budget. **Call before passing `attributes` to an ingest tool.** |
| `grill_docs_list` | List the documents in the project namespace (auto-pages internally). |
| `grill_projects` | List the projects your key can access (`product: "grill"` or `"primecut"`). Useful with account-level keys. |
| `grill_explain` | Returns a markdown explanation of how Grill works. No arguments, no auth — useful for self-documenting agents. |

> **`doc_id` vs `job_id`:** when an ingest reaches `done`, the document's `doc_id` in Grill equals the `job_id` returned by ingest. Keep that value to feed `doc_filter` on subsequent searches.

### `grill_ingest` and `grill_ingest_sync`

Both tools accept the same arguments. Provide **exactly one** of `file_path`, `file_base64`, or `url`.

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `file_path` | string | one-of | Path readable by the **MCP server process** (absolute or relative to its cwd). Best for large files; avoids giant JSON payloads. |
| `file_base64` | string | one-of | Standard base64 of the file bytes. Fine for small files. |
| `url` | string | one-of | Public URL — the **server** fetches the file. Works with the hosted endpoint (no laptop paths needed). |
| `filename` | string | no | Original basename (`report.pdf`). With `file_path`, defaults to the path basename; otherwise inferred from bytes. |
| `attributes` | object | no | Typed document attributes: a flat `name → value` map (string, number, boolean, or an array of one of those — an empty array needs its array type declared in `attribute_schema`; never `null`), e.g. `{"region":"emea","year":2025}`. **Call [`grill_attributes`](#grill-attributes) first** and reuse an existing name and type where one fits. Names match `^[a-z0-9_]{1,64}$` and are permanent per project. Sent as `X-Attributes` (max 64 names, 2048 chars of compact JSON — larger input is refused, never truncated). |
| `attribute_schema` | object | no | Type declarations for names in `attributes`, shaped `{"name": {"type": "…"}}`. Required for `encrypted_text` (always, except the reserved `labels`, which is typed `[]encrypted_text` automatically) and a new `datetime` name, and for the array type of an empty array; also use `float` / `[]float` to keep whole numbers as floats. A declaration that conflicts with the type already in force is rejected. Sent as `X-Attribute-Schema` (2048-char cap). |
| `labels` | object | no | **Legacy** — still accepted; prefer `attributes`. String key/value labels (no `:` or `,` in keys/values). The server translates each pair into a `"key:value"` string in the [`labels` typed attribute](/grill/concepts/ingestion#migrating-from-labels-meta-int). |
| `project_id` | string | no | Target project. Needed with account-level keys; defaults to `POMA_PROJECT_ID`. |
| `token` | string | no | API key — overrides `POMA_API_KEY` for this call. |

**Notes on `file_path`**

- Works only when the server runs on the **same machine** as the file. The hosted MCP at `mcp.poma-ai.com` cannot read laptop paths — use `file_base64` from there, or run the server locally.
- **`GRILL_INGEST_ALLOWED_PREFIX`** *(optional, env)* — when set, `file_path` must resolve under that directory after symlinks are evaluated. Non-regular files are rejected. Use this in any environment where untrusted prompts can reach the server.
- **`GRILL_INGEST_MAX_BYTES`** *(env, default 512 MiB)* — caps the payload size. Set to `0` to disable the limit (use with care).

**Picking `grill_ingest` vs `grill_ingest_sync`**

| Use case | Tool |
| --- | --- |
| Agent will search immediately and needs the doc indexed first | `grill_ingest_sync` |
| Long ingest; agent will check back later; conversation context is precious | `grill_ingest`, then poll with `grill_jobs_status` or reconnect with `grill_ingest_resume` |

### `grill_ingest_resume`

Reconnect to an already-running ingest job and wait for it to reach a terminal state. Use this when a previous `grill_ingest` returned a `job_id` (or the MCP connection dropped mid-ingest) and you don't want to re-upload the file.

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `job_id` | string | **yes** | Job ID from a prior ingest. |
| `token` | string | no | API key. |

Returns the same `job_id` + `events` payload as `grill_ingest_sync`.

### `grill_ingest_batch`

Upload up to 50 files in one tool call with controlled concurrency. Returns the `job_id`s as uploads complete — does **not** wait for indexing. Pair with `grill_jobs_status` to monitor.

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `file_paths` | array of string | **yes** | Up to 50 paths readable by the MCP server process. |
| `concurrency` | integer | no | Upload concurrency. Default `5`, max `10`. Use `1` on free-tier accounts. |
| `attributes` | object | no | Typed attributes applied to **every** file in the batch. Same rules as on `grill_ingest` — call `grill_attributes` first. |
| `attribute_schema` | object | no | Type declarations for `attributes`, as on `grill_ingest`. |
| `project_id` | string | no | Target project (account-level keys). Defaults to `POMA_PROJECT_ID`. |
| `token` | string | no | API key. |

Returns `{ results, submitted_count, failed_count, quota_exceeded_count }`. `quota_exceeded` entries (HTTP 403 from the queue) are retryable once running jobs finish.

### `grill_jobs_status`

Get status snapshots for up to 50 jobs in a single call. Non-streaming — call on an interval if you want progress.

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `job_ids` | array of string | **yes** | Up to 50 job IDs to query. |
| `token` | string | no | API key. |

Returns `{ results, pending_count, done_count, failed_count }`. Each result has `{ job_id, status, is_terminal, error? }`.

### `grill_search`

Hybrid search across the project namespace. Result count is bounded server-side by relevance and a token budget — there is **no `top_k`**.

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `query` | string | **yes** | Natural-language search query. |
| `doc_filter` | string | no | `doc_id` (= `job_id` from ingest) to restrict search to one document. |
| `exclude_doc_ids` | array of string | no | Doc IDs to exclude from results (max 100). Useful in agent loops to avoid re-citing docs already shown. |
| `return_assets` | boolean | no | Return cited docs' figures/tables in the response `assets` field, keyed by `doc_id` (images are base64 data URIs). |
| `attribute_filters` | array of `{attr, op, value}` | no | Filter on typed attributes; clauses AND together. Operators depend on the attribute's type (see [Retrieval](/grill/concepts/retrieval#filtering-on-typed-attributes)). A name the project has never used matches **nothing** — check names and types with `grill_attributes`. |
| `project_id` | string | no | Target project (account-level keys). Defaults to `POMA_PROJECT_ID`. |
| `token` | string | no | API key. |

Response (example):

```json
{ "context": "<context>This is what is relevant […] inside your document.</context>" }
```

The response also carries a `scope` object naming the project and namespace that were searched (surface it in agent output so users know which project answered), plus `assets` when `return_assets` is set. The `context` field is the main payload — drop it straight into your LLM prompt. See [RetrievalContext format](/grill/concepts/retrieval-context) for the wrapper grammar (`<doc>`, inline `[pN]` / `[…]` markers, sandwich ordering, citations).

### `grill_attributes`

List the typed attributes already declared in the project — each `name` with its `type` — plus `max_names`, the per-project cap on distinct names (64). Wraps [`GET /v3/grill/attributes`](/grill/reference/api#get-grill-attributes).

| Argument | Type | Required | Description |
| --- | --- | --- | --- |
| `project_id` | string | no | Target project (account-level keys). Defaults to `POMA_PROJECT_ID`. |
| `token` | string | no | API key. |

```json
{
  "attributes": [
    { "name": "region", "type": "string" },
    { "name": "year", "type": "int" }
  ],
  "max_names": 64,
  "note": "…",
  "scope": { "project_name": "…", "hint": "…" }
}
```

**Agents must call this before declaring a new attribute.** Attribute names and their types are permanent per project and count against `max_names`, so an agent that invents a fresh name per document (`year`, `doc_year`, `publication_year`) burns the budget and splits every later filter across synonyms. The rule:

1. Call `grill_attributes` before the first ingest that passes `attributes`, and before building `attribute_filters` for `grill_search`.
2. **Reuse** an existing name and type where one fits — send `2025` to an existing `year:int`, not `"2025"` and not a new `doc_year`.
3. Declare a new name only when nothing fits.
4. If the call fails with `retryable: true` (the API's `503` — the schema is temporarily unreadable), wait and retry. **Never treat a failed call as "no attributes exist"**, and don't declare new names until the list can be read.

Example flow:

```text
grill_attributes {}
  → attributes: [region:string, year:int]
grill_ingest_sync {
  "url": "https://example.com/q3-report.pdf",
  "attributes":       { "region": "emea", "year": 2025, "review_notes": "…" },
  "attribute_schema": { "review_notes": { "type": "encrypted_text" } }
}
  → region, year reused; review_notes is the one new name
grill_search {
  "query": "operating margin",
  "attribute_filters": [ { "attr": "year", "op": "gte", "value": 2024 } ]
}
```

If your client doesn't list `grill_attributes`, update the server (or the plugin) to the current release.

### `grill_explain`

Self-documentation tool. Takes no arguments and requires no authentication — returns a markdown explanation of how Grill works (ingest, search, result format, how to get an API key). Useful for agents that want to introspect their own capabilities, and for "explain this server" probes.

## Output shapes

`grill_ingest_sync` waits for a terminal status and returns events:

```json
{
  "job_id": "100c65a03a304aa343a1518aa79e8300-20260414T083549Z",
  "events": [
    { "status": "pending" },
    { "status": "chunking" },
    { "status": "chunked" },
    { "status": "grilled" },
    { "doc_id": "xxxxxx-xxxxx-xxxxx-xxxxx" }
  ]
}
```

`grill_ingest` (async variant) returns the same `job_id` immediately, without the trailing events. Ingest outputs also include a `scope` object naming the project/namespace the document landed in.

On error, the MCP response sets `isError: true` and returns a structured error:

```json
{ "error": "job failed: unsupported file type", "code": "job_failed", "retryable": false }
```

`code` is one of `missing_token`, `invalid_input`, `auth_expired`, `payment_required`, `project_protected`, `forbidden`, `too_many_jobs`, `upstream_error`, `transport_error`, `parse_error`, `job_failed`, `stream_error`. When `retryable` is `true`, `retry_after_seconds` may suggest a backoff.

## HTTP mode

Both Go and Node binaries support a long-lived HTTP server for custom integrations:

```bash
POMA_API_KEY=your-key poma-grill-mcp -http :8080            # Go
POMA_API_KEY=your-key node node/dist/index.js -http :8080   # Node
```

- **MCP endpoint:** `POST http://localhost:8080/` (standard streamable HTTP MCP).
- **Health check:** `GET http://localhost:8080/health`.

### Large uploads — `POST /ingest-upload` (Go binary only)

The Go binary exposes a dedicated upload endpoint that bypasses MCP framing for big files. Same auth as MCP — `x-api-key`, `Authorization: Bearer`, or fall back to `POMA_API_KEY` on the server.

- **Raw body:** send file bytes as the body. Pass the basename via the `?filename=…` query string or the `X-Filename` header.
- **Multipart:** `Content-Type: multipart/form-data` with a part named `file`.
- **Response:** `201 Created` with `{"job_id": "…"}`. Size limits follow `GRILL_INGEST_MAX_BYTES`.

```bash
curl -sS -X POST "http://localhost:8080/ingest-upload?filename=report.pdf" \
  -H "x-api-key: $POMA_API_KEY" \
  -H "Content-Type: application/octet-stream" \
  --data-binary @./report.pdf
```

The returned `job_id` is the same one you'd pass to `grill_search`'s `doc_filter` once indexing completes.

## Docker (Go only)

Images are tagged by **commit SHA** (no `latest` tag) — pick a tag from [the package registry](https://github.com/poma-ai/poma-grill-mcp/pkgs/container/poma-grill-mcp) first:

```bash
# HTTP mode
docker run -e POMA_API_KEY=your-key -p 8080:8080 \
  ghcr.io/poma-ai/poma-grill-mcp:<tag> -http :8080

# Stdio (default entrypoint)
docker run -i -e POMA_API_KEY=your-key ghcr.io/poma-ai/poma-grill-mcp:<tag>
```

A Node Docker image is not currently published — use `npx` or `npm install -g` instead.

## Flags and environment

| Flag | Default | Description |
| --- | --- | --- |
| `-input <path\|->` | — | Stdio mode: MCP on stdin (`-`) or a file path. |
| `-http <addr>` | — | HTTP mode (e.g. `:8080`). Mutually exclusive with `-input`. |

Both Go and Node accept the same flags.

| Env var | Default | Description |
| --- | --- | --- |
| `POMA_API_KEY` | — | API key used when no `token` is passed at call time — project key (`poma_prod_gr_…`) or account key (`poma_acc_…`, then also set a project). |
| `POMA_PROJECT_ID` | unset | Default project for account-level keys; per-call `project_id` overrides it. |
| `POMA_API_BASE_URL` | POMA prod | Override the API base for self-hosted/dev backends. |
| `POMA_STATUS_API_BASE_URL` | POMA prod | Override the status/SSE base for self-hosted/dev backends. |
| `GRILL_INGEST_ALLOWED_PREFIX` | unset | Restrict `file_path` to a directory tree (post-symlink). |
| `GRILL_INGEST_MAX_BYTES` | `512 MiB` | Max upload payload. `0` disables the limit. |

## Tips for agent-driven workflows

- **For large files, instruct the agent to use `file_path`.** Most clients default to `file_base64`. A short system-prompt cue like "When the user gives a local path, prefer `file_path` for `grill_ingest_sync`" saves enormous amounts of tokens.
- **Capture `job_id` after ingest.** The agent can then issue follow-up `grill_search` calls scoped to that document with `doc_filter` — this is the difference between "search my whole namespace" and "chat with this PDF" UX.
- **Combine ingest + search in one turn.** With `grill_ingest_sync` followed by `grill_search`, the agent can answer a question about a freshly uploaded file in a single response.
- **Check attributes before tagging.** Have the agent call `grill_attributes` before any ingest that sets `attributes`, reuse existing names, and retry (never assume "none") when the call fails.
- **Bulk ingest folders** with `grill_ingest_batch` + `grill_jobs_status` — submit everything once, poll status until terminal.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| `file_path: not found` | Hosted MCP is being asked to read a laptop path | Use `file_base64`, or run a local binary (Option A). |
| `payload too large` | File exceeds `GRILL_INGEST_MAX_BYTES` | Raise the env, use `/ingest-upload` (Go HTTP mode), or POST directly to the REST API. |
| `forbidden` / `401` on `grill_*` | Key belongs to a `primecut` project, or an account-level key was used without selecting a project | Use a Grill project key (`poma_prod_gr_…`), or pass `project_id` / set `POMA_PROJECT_ID` with your account key. `grill_projects` lists what the key can see. |
| `file_path` rejected with "outside allowed prefix" | `GRILL_INGEST_ALLOWED_PREFIX` is set | Move the file under the prefix, or unset the env if you trust the caller. |
| Agent looks for `primecut_*` tools and finds nothing | `poma-grill-mcp` only ships `grill_*` tools | Install [`poma-mcp`](/mcp/poma-mcp) alongside this server. |

## See also

- [`poma-mcp`](/mcp/poma-mcp) — PrimeCut MCP server (`primecut_*` tools).
- [Grill quickstart](/grill/getting-started/quickstart)
- [Grill API reference](/grill/reference/api)
- [RetrievalContext format](/grill/concepts/retrieval-context)