Agent Runtime
Plans and runs multi-step tasks, with a calibrated confidence score on every step.
- Plan
- Act
- Score
Dermexcel briefs, runs and reviews AI agents the way a good routine works: a few well-chosen tools, applied in order, with every step traced so you can see exactly what each one did.
curl https://api.dermexcel.dev/v2/runs \ -H "Authorization: Bearer $DX_KEY" \ -F "brief=@ticket.json" \ -F "routine=support"
Like any routine worth keeping, it runs in order: clean the context, do the work, protect the result. One endpoint, one event schema, every step logged.
Every brief is tidied before an agent sees it: duplicates merged, personal data redacted, context trimmed to what the task needs.
POST /v2/runs
The agent plans, calls your tools and checks its own work. Every step is traced and scored, so you can see what each one did.
planner.v3 · tools · self-check
Guardrails check every result against your policies. Anything uncertain goes to a person, a queue or a webhook instead of a customer.
guard.policies → human · webhook
Turn on what you need and leave the rest out. Every module speaks the same event schema, so an agent routine is a list, not an integration project.
Plans and runs multi-step tasks, with a calibrated confidence score on every step.
Decides which agent, model or person takes each task, using rules you version like code.
Keeps what your agents remember in step with your CRM, help desk and warehouse.
SQL access to runs, tool calls, costs and overrides, with scheduled exports.
Four steps, all of them yours to call directly. Nothing here is a black box you cannot log, replay or export.
Describe the task and hand over the tools it may use. Inputs are cleaned and personal data redacted on arrival.
The agent plans, calls tools and checks its own work, with every step traced and scored as it goes.
Guardrails check each result. Anything below your confidence threshold goes to a person, not a customer.
Every override becomes an eval. Ship a new routine behind a flag and compare it on live traffic.
Run in production by teams who page someone when it breaks.
Each routine is a template you can fork: the same modules with different tools, thresholds and retention.
Tickets in, resolved replies out, with anything uncertain handed to a person before a customer ever sees it.
Illustrative routine. Dermexcel is a fictional product built for this template.
"We replaced a support macro library two people maintained by hand with one routine. The part I did not expect was replaying a week of tickets against a new agent before switching."
"Sandbox key on Monday, first real tickets answered on Thursday. The SDK did what the docs said it would, which is rarer than it should be."
"Our security team cares where customer data lives. Running the agents in our own region, with a trace of every tool call we can export, is what got this approved."
Everything else is in the docs, and the docs are versioned with the API.
Project-scoped bearer keys, rotated from the dashboard or the API. Server-side keys can be pinned to an IP range; browser SDKs get short-lived tokens minted by your backend.
60 requests per second per project on the standard plan, burst to 200. Every response carries the remaining budget, and a 429 tells you when to retry rather than making you guess.
In the region you pick, and only there. Conversations can be dropped the moment a run finishes, or kept for a retention window you set per routine. Every read is in the audit log.
Agent workers ship as a container you can run in your own VPC, with the control plane still managed. Model and planner updates arrive as signed images you choose to promote.
99.95% on the API and a p95 latency target per region, measured from our edge, with credits applied automatically from the same numbers on the status page.