Back-office agents
Match invoices to purchase orders, chase missing documents, reconcile statements and post to the ERP with exceptions routed to a person.
Multi-step agents that read documents, check systems of record, decide what to do and do it, with a human approval step wherever the stakes call for one.
Most agent demos fail in production for the same reasons: they cannot see the systems the work lives in, nobody can tell why they did something, and there is no safe way to let them act. We build agents the other way around, starting from the workflow, the permissions and the audit trail.
Each agent is a narrow worker with named tools, a written policy and an evaluation set. It runs on open-weight models inside your network, so the documents it reads and the actions it takes never pass through a vendor's servers.
| Models | Llama, Qwen, Mistral, DeepSeek or gpt-oss, chosen by bake-off |
|---|---|
| Orchestration | Typed tool calls, queued jobs, retries and idempotent actions |
| Human in the loop | Per-action approval rules, by amount, risk or customer |
| Audit | Every prompt, tool call and result logged with the user and case |
| Deployment | On-premises, private cloud or air-gapped |
Match invoices to purchase orders, chase missing documents, reconcile statements and post to the ERP with exceptions routed to a person.
Read inbound email, forms and attachments; classify, extract fields, open the right ticket or case and draft the first reply.
Gather facts across internal drives, databases and approved sites, then write a sourced brief in your house format.
Watch queues and dashboards, explain what changed, and propose the next action for an operator to approve.
An AI agent is a model that can take steps. It reads an input, checks one or more systems, decides what to do and calls a tool to do it. Custom AI agent development means building that worker for one defined job in your business, with the tools, rules and approval steps that job requires.
LLM.co builds agents as part of its custom AI development practice. Each agent runs on open-weight models as private AI inside your network. The invoices, emails and records it handles stay on infrastructure you control.
Agents fit work that follows a pattern but needs judgment at a few points. Good candidates have a clear trigger, a defined set of systems, and an outcome someone can check. Accounts payable matching, intake triage and document chasing are typical examples.
Agents are a poor fit where a fixed rule already works. If a script or an RPA flow handles the case reliably, keep it. We also advise against agents for one-off tasks with no history to test against, because there is nothing to build an evaluation set from. Custom AI agent development at LLM.co starts with that check. We confirm the workflow has enough past cases to score against before designing anything.
We keep each agent narrow. It gets a written policy, a short list of named tools, and a service account with the least privilege that works. Tool calls are typed and validated before they run. Writes are idempotent, so a retry never posts the same entry twice.
Longer jobs run as queued tasks with retries and timeouts. When an agent is unsure, or a case falls outside its policy, it stops and hands the case to a person with a summary of what it found. Several narrow agents usually beat one broad one, because each can be scored and approved on its own.
The main risks with agentic AI are wrong actions, unexplained actions and actions taken with too much access. Each one has a specific control.
You receive the agent source code, prompts and policies, tool adapters, the evaluation set and its reports, and the deployment code for your environment. You also get runbooks covering approval settings, the kill switch and model upgrades. There is no platform license, and you can extend or retire the agent without us.
Write down every step, system and decision a person makes today, and mark which ones an agent may take alone.
Expose only the APIs and queries the job needs, each scoped to a service account with the least privilege that works.
Build an evaluation set from real past cases and measure accuracy, cost and time per case before anyone relies on it.
Approval steps, rate limits, a kill switch and a full action log in your SIEM from day one.
It is the design and build of an AI agent for one defined job in your business. The agent reads inputs, checks your systems, decides on an action and carries it out through approved tools. It ships with a written policy, an evaluation set, approval rules and a full action log, and runs on infrastructure you control.
Only where you decide it may. Each action type gets its own rule: always ask, ask above a threshold, or act and report. Most teams start with approval on everything and relax it as the evaluation scores hold up. You can tighten the rules again at any time.
Anything with an API, a database or a structured export. That includes ERPs, CRMs, ticketing tools, email, document stores and data warehouses. Where there is no API, we prefer a vendor export or database view over screen automation, because screen automation breaks when the vendor changes the interface.
A two-week discovery sprint maps the workflow and confirms the case for an agent. A focused first agent usually reaches production in eight to twelve weeks. The Prototype phase takes three to four weeks, Harden four to six, and Deploy one to two. Integrations and approval design are the main variables.
Discovery is a fixed fee, and the build is quoted per phase after that. Cost depends on how many systems the agent reads and writes, how clean the data is, and how detailed the approval and audit rules need to be. Running costs depend on whether the models serve from your cloud account or your own hardware.
RPA follows fixed steps and fails when the input changes. An AI agent can read unstructured inputs such as emails and scanned documents and decide which step comes next. The two work well together. We often keep existing RPA for stable steps and add an agent where judgment is needed.
The agent runs as private AI on infrastructure you control, so documents and actions never pass through a vendor's servers. It uses scoped service accounts, and tool permissions are enforced outside the model. Every action is logged, and the design supports your existing SOC 2, HIPAA or similar control program.
Yes. Every agent we build runs on open-weight models served on hardware you own, in your own cloud account, or on an air-gapped network. Nothing calls a third-party model API unless you decide it may. The same agent can move from cloud to on-premises later with little change.
Tell us the workflow and where the data lives. An engineer, not a salesperson, replies within one business day with a first take on architecture and cost.