# Agent API — letting an external AI drive an account

`/api/v1/*` is a Bearer-key REST API that gives an external "AI agent"
(a ChatGPT custom GPT Action, another assistant, a script) the same
account-scoped abilities the dashboard SPA has: create and manage phone
agents, SIP trunks/accounts and numbers, read call history/transcripts/
recordings, and handle billing (check balance, start a top-up).

It is a thin auth layer on top of the existing dashboard API
(`App\Http\Controllers\Admin\ApiController`) — same validation, same
account scoping — reachable without a browser session.

## Getting a key

1. Log in and open <https://www.callagent.pro/admin/settings> → tab
   **API & Webhooks**.
2. Click **New key**, give it a name (e.g. "ChatGPT").
3. Copy the key shown (`cak_live_...`) — it is shown once, like a GitHub
   token. Hand it to the AI agent.
4. Revoke it any time from the same screen if you want to cut the AI off
   without touching anything else on the account.

(Endpoints behind that screen: `GET/POST /admin/api/agent-keys`,
`DELETE /admin/api/agent-keys/{id}` — session-authenticated, dashboard-only.)

## Using it

```
Authorization: Bearer cak_live_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
Content-Type: application/json
```

Full machine-readable description (importable as a ChatGPT Action, or into
any OpenAPI-aware tool): **`/openapi/agent-api.yaml`**. A short orientation
file for assistants lives at **`/llms.txt`**, and the landing page advertises
both via `<link rel="service-desc">`, so an agent that is merely pointed at
callagent.pro can find all of this without being told where to look.

That key is the *only* thing you need from a human. Every dashboard screen has
an endpoint here; driving the dashboard with a browser extension is slower,
breaks on any layout change, and asks the user for a password instead of a
revocable key.

### Typical flow: "create me a phone agent for my business"

```
GET  /api/v1/meta                 -> languages, voices, ASR models, LLM servers
POST /api/v1/agents               -> { name, prompt, primary_language_id, default_voice_id, ... }
POST /api/v1/agents/{id}/test-chat -> sanity-check the prompt with one text turn
POST /api/v1/agents/{id}/test-call -> { phone_number } — a real call to verify it end-to-end
PUT  /api/v1/numbers/{id}         -> point an existing number at the new agent
```

### Benchmarking the agent it just built

`POST /api/v1/agents/{id}/test-chat` with a `messages` array runs a whole
scripted conversation through the real pipeline (flow engine, tools,
knowledge base — only STT/TTS are swapped out, same harness as
`textchat.py`/`llmbench.py`) in a single request and scores every turn:

```json
POST /api/v1/agents/42/test-chat
{ "messages": ["Hi, what's the price of gold?", "And the weather in Berlin?"] }
```

```json
{
  "ended": true,
  "transcript": [ { "role": "caller", "text": "..." }, { "role": "agent", "text": "..." } ],
  "turns": [
    { "caller": "Hi, what's the price of gold?", "first_word_s": 1.9, "total_s": 3.1, "utterances": 2 }
  ],
  "timing_summary": {
    "turns": 2,
    "first_word_s": { "min": 1.7, "max": 1.9, "avg": 1.8 },
    "turn_total_s": { "min": 2.4, "max": 3.1, "avg": 2.75 }
  }
}
```

`first_word_s` is what a caller would perceive as time-to-first-spoken-word
— the number to watch when judging whether a prompt/flow/model combination
is fast enough for a live phone call. Rate-limited tighter
(`agent-api-heavy`, 6/min) since each call runs one or more real LLM turns.

### Typical flow: "get me logs / transcripts / recordings"

```
GET /api/v1/calls              -> recent call history
GET /api/v1/calls/{id}         -> transcript, captured variables, status,
                                   and audio_url (streams with the same Bearer key)
GET /api/v1/live               -> calls in progress right now
```

### Real-time call control

`DELETE /api/v1/live/{id}` hangs up a call in progress right now — the
intervention a script watching `GET /api/v1/live` has no other way to make
(a fraud pattern detected mid-call, an emergency). Same Asterisk AMI path as
the dashboard's own "End call" button. Transfer and mute are not available
yet.

### Pagination, filtering, bulk operations

`GET /api/v1/agents`, `/calls` and `/contacts` accept `limit` (default 50,
max 200) and `offset`, and return `meta: {total, limit, offset}` alongside
the list. `/calls` also accepts `status`, `date_from`, `date_to`, `search`;
`/agents` accepts `status`, `search`; `/contacts` accepts `search`.

`POST /api/v1/agents/bulk-delete` and `POST /api/v1/contacts/bulk-delete`
each take `{"ids": [...]}` (up to 500 / 5000) and return
`{ok, deleted, failed_ids}` — one request instead of one per row.

### Errors

Every non-2xx response, on every endpoint, is `{"error": "<slug>",
"message": "<text>"[, "errors": {...}]}` — `errors` (Laravel's native
per-field shape) only appears on 422s. See the `Error` schema and the
`Unauthorized`/`Forbidden`/`NotFound`/`ValidationError`/`RateLimited`
response components in the OpenAPI spec.

### The post-call webhook

`Agent.webhook_url`, once set, gets a JSON `POST` after every completed
call — see the `WebhookPayload` schema in the OpenAPI spec for the exact
shape (`datetime`, `phone`, `duration`, `summary`, `variables`), retried 3x
with backoff. This is a webhook YOUR server receives, not an endpoint on
this API.

### Typical flow: billing

```
GET  /api/v1/billing            -> balance, plan, this month's usage, invoices
POST /api/v1/billing/topup      -> { amount } (EUR, 5-500, step of 5)
                                    -> { approval_url } — hand this link to the
                                       human; PayPal payment can't be completed
                                       headlessly, only started
```

### Paying without a human at all (x402)

PayPal always needs someone to click. The x402 endpoints do not: they speak
HTTP 402 and settle in USDC on Base *inside* the request, so by the time you
read the response the balance has already moved.

```
GET  /api/v1/billing/x402       -> balance, can_call, what we accept, and the
                                   last ten settled payments. Free to call and
                                   never challenged — an agent has to be able
                                   to learn it is short before it can decide
                                   to pay. `accepts` is null if x402 is off.
POST /api/v1/billing/x402/topup/{amount}
```

The handshake: call the top-up once with no payment header and it answers
**402** with a `PAYMENT-REQUIRED` header. Sign an EIP-3009
`TransferWithAuthorization` for that amount, retry the identical request with
`PAYMENT-SIGNATURE`, and you get a 200 with the credit applied and a
`PAYMENT-RESPONSE` header carrying the settlement. Most x402 client libraries
do all of this for you.

`{amount}` must be one of `accepts.denominations_eur` — the price comes from
the path, never from the body, so there is nothing for a caller to inflate.
Follow `accepts.top_up_url` rather than assembling the URL yourself.

The response reports two currencies on purpose: `paid`/`paid_currency` is what
left the wallet (USDC), `credited`/`credited_currency` is what the balance
gained (EUR). Crediting is idempotent on `transaction`, so a 500 that says
`settlement_not_credited` must **not** be retried — that pays twice.

`POST /api/v1/agents` accepts the same challenge, so an agent can pay for the
phone agent it is creating, in the request that creates it.

### Worked example: a CRM lead qualifier

Outbound. The point is that the result is structured data — each answer the
caller gives comes back as a captured variable, so nobody has to listen to a
recording to find out what happened.

```
GET  /api/v1/meta                    -> language, voice, ASR model, LLM server
POST /api/v1/agents                  -> prompt = the qualifying questions.
                                        Set webhook_url so each result is
                                        pushed as the call ends.
POST /api/v1/contact-lists           -> the list a campaign draws from
POST /api/v1/contacts                -> the leads
GET  /api/v1/do-not-call             -> check BEFORE dialling. Reading this
                                        list is the whole reason it exists.
POST /api/v1/agents/{id}/test-chat   -> score the script on a scripted
                                        conversation, before a human hears it
POST /api/v1/agents/{id}/test-call   -> one real call
GET  /api/v1/calls?date_from=...     -> then /api/v1/calls/{id} per call
```

`GET /calls/{id}` returns `vars` as `[[key, value], ...]` — what the agent
captured — plus `summary`, the full `transcript` and `audio_url`. The post-call
webhook carries the same variables, which is usually better than polling.

### Worked example: a practice appointment desk

Inbound, against a real calendar. The agent offers only slots that are
genuinely free, books while the caller is on the line, and emails the
invitation.

```
GET  /api/v1/profile                 -> check google_calendar_connected FIRST
POST /api/v1/agents                  -> tools: "cal", or "secretary" for the
                                        whole bundle
GET  /api/v1/numbers
PATCH /api/v1/numbers/{id}           -> point the practice number at the agent
POST /api/v1/agents/{id}/test-call   -> verify end to end
GET  /api/v1/calls/{id}              -> transcript, and what got booked
```

**The one step that is not headless.** Connecting Google Calendar is an OAuth
consent and has to be done once by the human, in the dashboard. The API can
*read* whether it happened but cannot do it: `Profile.google_calendar_connected`
is read-only. Check it before building a booking agent — a `cal` tool with no
calendar behind it will happily take the call and then fail to book, which is
a worse outcome than refusing up front. If it is false, ask the user to connect
it and carry on once they have.

Two details worth knowing about `tools`:

- It is a comma-separated string of **built-in capability names**, not an
  OpenAI function-calling schema. `cal` gives the agent `check_availability`
  and `book_appointment`; `secretary` bundles contacts, calendar, scheduled
  calls, do-not-call, notes and email.
- The calendar actions are **create and delete**. There is no in-place edit, so
  a reschedule is a delete followed by a create — which is what the agent does
  on a call, and what you should expect to see in the transcript.

Contacts the agent looks up mid-call are the same records you manage through
`/api/v1/contacts`, so populating them over the API is what makes caller
recognition work.

## Scope, deliberately

Not exposed on `/api/v1/*`: superadmin, password change, campaign
management, DNC bulk tooling, real-time transfer/mute on a live call. An AI
agent managing "my phone agent" has no business with most of those, and
keeping the surface narrow keeps a leaked key's blast radius small. (Bulk
delete IS exposed for agents/contacts — see above — since "manage a
hundred contacts" is a legitimate scripted use case a human clicking
through the dashboard never hits.) Ask if a legitimate use case needs one
of the excluded ones opened up.

## Connecting a real phone line — Fritz!Box / DSL

"Configure it with my Fritz!DSL to work now" means routing calls between
the customer's existing Fritz!Box and a callagent.pro agent. There are two
different setups depending on what the customer wants to keep:

### A. Dedicate the line entirely to the AI (no phones behind the Fritz!Box anymore)

Take the SIP credentials the ISP already put on the Fritz!Box and register
*our* Asterisk with them directly — the Fritz!Box becomes unnecessary for
calls (it can stay for the physical DSL/router function only).

1. On the Fritz!Box admin UI (`http://fritz.box`) → **Telephony → Own
   Numbers → Edit** the DSL/internet number: the SIP registrar, username
   and password for the line are shown there (this is what "Fritz!dsl"
   means in practice — an AVM-branded SIP trunk).
2. `POST /api/v1/trunks` with `provider: "custom"`, `host`/`port` from that
   screen, `username`/`secret` copied from it, `direction: "inbound"` (or
   `"both"` if the account will also place outbound calls), and
   `inbound_assistant_id` set to the new agent's id.
3. Once `is_registered` comes back true (`GET /api/v1/trunks` — or trigger
   `POST /api/v1/trunks/{id}/retry` if it doesn't within a minute), incoming
   calls to that DSL number ring the AI agent directly.

### B. Keep the Fritz!Box's own phones, add the AI as one more extension (verified click-path, Fritz!OS on a 7530)

Useful when the customer wants some calls to still ring an office phone and
others (an overflow number, an after-hours divert) to reach the AI, or as
the easiest way to stand the integration up without touching the real line
at all.

1. `POST /api/v1/sip-accounts` to create a SIP extension on callagent.pro —
   pick a `username` (e.g. an unused-looking number like `100049`) and a
   `password`.
2. On the Fritz!Box admin UI (`http://fritz.box`, login with the box's
   password) → **Telephony → Phone Numbers → New Phone Number**.
3. Check **"Use internet phone number"**, set **Telephony provider** to
   **"Other provider"**. A row appears asking for:
   - **Phone number for registration** and **Internal phone number in the
     FRITZ!Box** — both: the SIP account's `username` from step 1.
   - **Display name** — anything (e.g. the agent's name).
4. Click through to the account-information screen and fill in:
   - **Username** and **Authentication name** — the SIP account's
     `username` (both fields take the same value).
   - **Password** — the SIP account's `password`.
   - **Registrar** and **Proxy server** — `callagent.pro` (both the same
     host; leave **STUN server** blank).
5. Save. Back on **Telephony → Phone Numbers**, the new number shows up
   in the list with a green status dot next to the existing PSTN/VoIP
   lines (a 1&1 or Telekom DID, say) once it registers —
   `GET /api/v1/sip-accounts` (or the trunk's registration event log) can
   confirm `is_registered: true` from our side too.
6. To make the AI answer the **existing business number(s)** rather than
   just this new one: **Telephony → Call Handling → Call Diversion → New
   Call Diversion**. For each existing DID that should reach the AI:
   - **For calls from/to**: the existing number (e.g. `23573844`).
   - **Diversion to**: the *same* number (this isn't a forward to a
     different external number — it's telling the box which of its
     internet lines to send the call out over).
   - **Via Telephone Number**: the new internet number from step 3
     (`100049`) — this is the field that actually routes the call to
     callagent.pro instead of ringing a handset.
   - **Condition of diversion**: `immediately` (or restrict by schedule
     for an after-hours/overflow-only setup — the box also supports
     enabling/disabling the whole diversion table on a schedule).
   - Repeat one diversion row per DID that should reach the AI; leave any
     number without a row ringing office phones as before.
   `GET /api/v1/calls` / `/live` on our side will then show inbound calls
   arriving with that DID as the called number, right after the first test
   call.

   (Also visible on the number's own "Additional Settings" screen: caller
   ID passthrough, and SRTP media encryption the box already advertises by
   default — no extra configuration needed for that.)

### A. Dedicate the line entirely to the AI

If the customer's number is itself already a SIP line rented from AVM/the
ISP (branded "Fritz!DSL"), its own registration credentials are visible on
the same **Telephony → Phone Numbers → Edit** screen for that number
(Registrar/Proxy/Username/Password — the provider is shown as the ISP's
name instead of "callagent.pro"). Copying those into
`POST /api/v1/trunks` (`provider: "custom"`, same field mapping as step 4
above) registers *our* Asterisk directly as that line's SIP client instead
of the Fritz!Box, so the AI receives its calls with no diversion step and
no Fritz!Box in the path at all. Only do this once B is proven to work —
it's harder to undo if something's misconfigured, since the Fritz!Box stops
handling that number's calls entirely.
