Hire an AI employee that clocks in and never clocks out
A Digital FTE is not a chatbot. It is a scoped role that reads your inboxes, your ERP, and your ticket queue, does the repeatable part of a job end to end, and escalates the rest with its reasoning attached. Built spec-first, so you approve what it will and won't do before it touches anything.
- Shift coverage, no rota
- 24/7
- Years building production systems
- 3+
- Projects shipped end to end
- 12+
Shift coverage, no rota
Years building production systems
Projects shipped end to end


Built on
- OpenAI Agents SDK
- LangGraph
- General Agents
- n8n
- Model Context Protocol
- TypeScript
- Python
- Postgres + pgvector

Where the payroll hours actually go
Your team is not slow. The handoffs are.
Every operations team I have looked at loses the same hours in the same four places, and none of them are the part of the job you hired that person for.
The work stalls between systems, not inside them
CRM, ERP, inbox, and spreadsheet each do their job fine. A person is the glue: copying a value, checking a policy, chasing a status. That cost never shows up on any tool's invoice.
Rule-based automation dies on the first exception
A recorded script breaks the moment a field moves or a supplier sends a different format, so someone ends up reviewing the automation as well as doing the work.
A chatbot answers; it does not finish the job
Explaining the refund policy is not issuing the refund. Deflection numbers look good while the queue length stays exactly where it was.
Nobody signs off on work they cannot audit
The blocker on most agent projects is not accuracy. It is accountability. Without scoped permissions, a log, and a human checkpoint on the risky steps, the rollout stops at the security review.
What a Digital FTE actually is
A role you scope and staff, not a tool you have to drive
The unit of delivery here is a job description, not a feature list. We write down the role, the systems it may touch, the decisions it may make alone, and the ones that stop for a human. Then it runs that role on your infrastructure.
A chatbot
Answers a question and hands the work back to a person.
Never writes to your systems of record.
An RPA script
Replays a fixed sequence of clicks against a fixed screen.
Cannot reason about an exception it has not seen before.
A Digital FTE
Owns a scoped role: reads the situation, decides, acts through permissioned tools, and reports what it did.
Still escalates anything outside its written boundary.
Inside one shift
Observe, decide, act, report, then round again
An agent is not one call to a model. It is a loop: it reads the whole situation, writes down the next move and why, executes against your real systems, then records what actually happened so the next pass is better informed.

Stage 01
Observe
Read the whole situation
Pulls the record from every system that matters: the invoice, the PO, the contract terms, the prior tickets. It does not reason from a single message.
Stage 02
Decide
Write down the next move
Chooses the next action and states the reason in plain language. That reasoning is stored, so a disputed decision can be read back rather than guessed at.
Stage 03
Act
Call the real tools
Executes through a permissioned tool layer wired into your ERP, CRM, ticketing, or inbox. Anything above the threshold you set stops at an approval gate with a named owner.
Stage 04
Report
Close the loop out loud
Logs the outcome, updates the record, and files the exception queue, so the morning briefing is a summary of work done rather than a request for instructions.
- obsinvoice INV-88301 · PO-4471 · goods receipt GR-9920
- thinkline 3 over-billed by 4.2% → outside 2% tolerance
- toolerp.hold_payment(inv_88301) → ok
- gate>$25k requires approval → queued to finance lead
- logexception filed with full reasoning attached
The loop exits on one of two conditions: the goal is met, or the agent hits a boundary you defined and escalates to a named human with its full reasoning attached. It does not guess its way past a wall.
How it is actually engineered
Loop, harness, and graph engineering: the three layers that decide reliability
Picking a good model is the easy part, and it is not where agents fail. What separates a demo from something an operations lead will sign off on is the engineering around the model: how often it is allowed to think again, what it is handed and what it may touch, and whether its control flow is written down or improvised. These are the three layers I build every Digital FTE on.


Loop Engineering
When it thinks again, and when it stops
The loop is the agent's heartbeat: observe, decide, act, then re-read the result and decide whether it is done. The engineering is in the exits, not the iterations: a step budget, a token and cost ceiling, a convergence check so it does not re-run the same call with the same input, idempotency keys so a retry cannot double-post an invoice, and a hard stop that escalates instead of guessing.
- Step, token, and cost ceilings enforced per run
- Idempotency keys on every write so retries are safe
- Repeat-call detection instead of a silent infinite loop
- Termination on goal met, boundary hit, or budget spent, never on luck

Harness Engineering
What surrounds the model call
The harness is everything that is not the model: which records get assembled into context, what the tool schemas look like, which credentials the call runs under, how the output is validated before it is trusted, and what gets written to the trace. Most production failures I have debugged were harness failures (a stale record, a loose tool contract, an unvalidated field) rather than reasoning failures.
- Context assembled from systems of record, not pasted history
- Typed tool schemas with validation on the way in and the way out
- Scoped credentials per agent, checked at the tool layer
- Prompt-injection filtering on anything arriving from outside the org
- Every prompt, call, and outcome traced and replayable

Graph Engineering
Control flow you can read on a page
Past a certain complexity, a single loop stops being honest about what the agent does. The work becomes a graph: nodes are steps, edges are conditions, and the state is checkpointed at each hop, so a run can pause at an approval gate on Friday and resume on Monday exactly where it stopped. It also makes review possible: the diagram is the behaviour, and a failed node has a deterministic fallback rather than a retry loop.
- LangGraph state machines instead of one open-ended prompt
- Checkpointed state, so runs survive a restart or an approval wait
- Human approval gates modelled as nodes, not bolted on afterwards
- Deterministic fallback branch on every step that can fail
- Supervisor plus specialists when one role needs several skills
The three are one system. A tight loop with no harness burns budget quickly and confidently. A strong harness with no graph works right up to the first exception that needs two steps and a human. The reason a Digital FTE can be handed to your team is that all three are written down before the build starts.
Roles you can staff
Seven jobs a Digital FTE can hold down
Each of these is scoped as a role with a written boundary, not a feature bolted onto an existing tool. Most engagements start with one and add the next once it is running unattended.

Operations analyst
The Monday-morning role. It reconciles the week across your systems, audits what moved against what should have moved, and lands a briefing in your inbox before anyone opens a laptop, with the exceptions ranked and the reasoning attached.
- Cross-system reconciliation on a schedule
- Ranked exception queue instead of a raw report
- CEO briefing generated from live records, not a template

Finance & invoice clerk
Three-way matching between invoice, purchase order, and goods receipt, with the clean cases going straight through and only genuine discrepancies reaching a person. Approvals above your threshold still stop for a named owner.
- Invoice, PO, and receipt matched automatically
- Discrepancies routed with their evidence
- Threshold-based human approval, never a silent auto-approve

Inbox & comms handler
Watches Gmail, WhatsApp, and the shared inboxes; classifies what arrives, drafts the reply in your voice against your own policies, files the attachment where it belongs, and flags anything that needs a human before it goes out.
- Multi-channel monitoring (Gmail, WhatsApp, shared inboxes)
- Drafts grounded in your own policy documents
- Send-gate on anything customer-facing you mark sensitive

Customer support agent
Grounded in your help centre and past tickets, so it resolves the routine request end to end by issuing the change, not just explaining it, and escalates the rest with the full context already attached.
- Retrieval over your real help centre and ticket history
- Executes the change through scoped API permissions
- Escalations arrive with context, not a transcript

Back-office data operator
The copy-paste role nobody wants: pulling values between Oracle, Odoo, spreadsheets, and your CRM, validating them against the source, and keeping both sides in step without a nightly export ritual.
- ERP and CRM write-backs with idempotent retries
- Validation against the system of record before write
- Sync failures surfaced as alerts, never swallowed

Task & project coordinator
Ingests incoming work, validates it, and routes it to the right owner with a deadline. The intake-to-assignment pipeline runs as an event stream instead of a person triaging a queue.
- Event-driven intake with automatic validation
- Routing rules you can read and change
- Auto-generated task lists and follow-up chasing

Multi-agent operations team
When one role is not enough: a supervisor that decomposes the job, specialists that own each step, and shared memory so context survives the handoff. It is the same structure you would use if you were hiring three people instead of one.
- Supervisor plus specialist topology
- Persistent memory across steps and sessions
- Deterministic fallback when a specialist fails
Not sure which role should go autonomous first?
Bring the process you want off your team's desk. You leave the call knowing what it would cost to run, what it would take to build, and whether an agent is even the right answer.

Shipped work
Systems already running this way

Digital FTE Dashboard
A real-time analytics dashboard that consolidates Digital FTE performance, agent activity, and approvals across multiple clients, with Oracle and Odoo data in one place.

Task Automation Engine
An AI task engine that ingests, validates, and routes incoming work items automatically, using Redpanda event streaming with auto-generated workflows.

Odoo Integration Hub
A custom integration hub that connects Odoo ERP to external services with automated invoicing, catalog and recipe management, and two-way data sync.
Built for what actually gets approved
The parts nobody demos, that decide whether it ships
Model quality is rarely what stops an agent reaching production. These are the layers that get it through a security review.
Capability
What it means for you
Tool and API calling
Typed, permissioned calls into your CRM, ERP, database, and third-party tools. Never screen scraping.
Human-in-the-loop gates
Manual review checkpoints built into the high-stakes steps, with a named owner rather than a shared inbox.
Role-based access control
Each agent reaches only the data and actions it was explicitly authorised for, even if it is asked otherwise.
Audit trail and logging
A replayable record of every decision, tool call, prompt, and outcome, retained on your side.
Persistent memory
State and context that survive across runs, so an agent picks up a long-running task where it left off.
Fallback behaviour
When it hits a case it was not designed for, it stops and routes to a person with the trace attached. It does not improvise.
Local-first deployment
Cloud, on-premise, or hybrid, including open-weight models when the data genuinely cannot leave your network.
Evaluation before rollout
Every change is measured against a fixed test set built from your real cases before it reaches production.
Headcount maths
The honest comparison, including where a person still wins
A Digital FTE is not a replacement for judgement, relationships, or accountability. It is a replacement for the repeatable half of a role, and it is worth being precise about which half.
| Human hire | Digital FTE | |
|---|---|---|
| Hours | 9–5, weekends off | 24 / 7 |
| Ramp-up | Weeks of onboarding | Deployed with your SOPs |
| Cost | Salary + benefits + overhead | Fixed subscription, no overhead |
| Scaling | Hire again | Clone the agent |
| Consistency | Varies by day | Same output, every run |
| Supervision | Needs active management | Self-operating, exception-only |
Keep this with a person
- Judgement calls with no precedent to reason from
- Relationships, negotiation, and anything reputational
- Owning the outcome when something goes badly wrong
Hand this to the agent
- The same decision made a hundred times, identically
- Work that arrives at 2am, on a Sunday, in peak season
- Cross-system checking nobody has ever enjoyed doing
Industries
Where the first Digital FTE usually lands
The pattern repeats across sectors: high-volume repeatable work, a system of record nobody wants to touch twice, and a person acting as the bridge between two tools.
- E-commerce & retail ops
- Logistics & supply chain
- Fintech & finance ops
- Insurance back office
- Healthcare administration
- Real estate
- Legal ops
- Professional services
- Manufacturing
- SaaS customer operations
- Education administration
- Agencies & studios


Tech stack
The runtime behind the role
Chosen for boring reasons: everything here is something I have run in production and can hand over without a translation layer.
Agent runtime
OpenAI Agents SDK · LangGraph · General Agents · MCP
Reasoning models
GPT family · Claude · Gemini · Open-weight for private deploys
Memory & retrieval
PostgreSQL · pgvector · Qdrant · Redis
Integrations
Oracle · Odoo · Gmail / WhatsApp · REST & webhooks · n8n
Application
Next.js · TypeScript · Python · FastAPI
Operations
Docker · Vercel · GitHub Actions · Structured tracing
How the build runs
Seven steps from role audit to handover
Every stage has an exit you can take if the numbers stop making sense. Nothing is built before you have approved what it is for.
Role audit
We map the job as it is really performed, including the shortcuts nobody documented, and score each step by volume, cost, and risk to find where an agent earns its keep.
Written spec
A job description for the agent: scope, systems, decisions it owns, decisions it escalates, and what it will refuse to do. You approve it before any code exists.
Evaluation set
A scored set of your real cases, agreed up front, so quality is a number that moves rather than an opinion argued about later.
Prototype on real data
A working agent runs against your sample data early, so you see its actual behaviour, and its actual failure modes, while changing course is still cheap.
Integration & guardrails
Scoped credentials, tool layer, approval gates, fallback paths, and tracing go in before the pilot, not after someone asks for them.
Supervised pilot
The agent runs beside your existing process on live traffic. You watch its decisions, its cost, and its edge cases with nothing at stake operationally.
Rollout & handover
Staged cutover with monitoring, cost ceilings, and a documented rollback. Then the codebase, prompts, and runbooks are handed to your team. You own all of it.
Control
Autonomy your security review will actually approve
Autonomy with a hand on the brake
Every Digital FTE ships with the controls that make an operations lead comfortable signing off on it.
- Explicit permission boundaries defined before deployment, not discovered after
- High-stakes actions pause for human approval with the reasoning attached
- Every prompt, tool call, and outcome logged and replayable
- Scoped credentials per agent, so it gets only the access its role requires
- Prompt-injection filtering on anything arriving from outside your org
- Local-first and on-premise options when data cannot leave your network
- Cost ceilings and rate limits so a runaway loop is capped, not invoiced
- Your data stays in your environment and is never used to train models

Spec before code, every time
You approve a written role definition first, so the argument about what was promised never happens on delivery day.
Built for production, not for a demo
Integrated with your real systems from the first sprint. A prototype that only works on curated sample data is not a milestone.
You own the whole thing
Codebase, prompts, evaluation set, and infrastructure access hand over on delivery. No vendor lock, no black box.
I will tell you when it is the wrong answer
If a scheduled job, a better form, or a smaller script solves your problem, you hear that on the first call rather than after the invoice.
Engagements
Three ways to put one on the payroll
Starter Agent
$2,500/project
- Single workflow
- Up to 2 integrations
- Basic monitoring
- 2-week delivery
Production Agent
$6,000/project
- Multi-step workflows
- Up to 5 integrations
- Full monitoring & alerts
- Human-in-the-loop approvals
- 4-week delivery
Agent Retainer
$1,500/month
- Ongoing operation
- Continuous tuning
- Priority support
- Monthly reporting