AI Agent Engineer
AI Agent Engineer
XtendOps- Posted 8 hours ago
- Be among the first 10 applicants
Job Description
About the role
XtendOps builds AI agents that handle real work for enterprise clients — reading incoming requests, deciding what to do, calling tools across the client's systems, and either acting or preparing work for a human to approve. These agents run against live customer traffic every day.
You will own one or more of them outright: design the flow, build the tools, wire the integrations, ship it, and keep improving it. This is a backend engineering role — we build systems around models, we don't train them.
Key Responsibilities
- Build and ship AI agents end to end, from design through production
- Design and implement the tools and MCP servers agents call, including schemas and descriptions that models use correctly
- Integrate third-party APIs and internal services behind those tools — auth, retries, rate limits, idempotency
- Write and iterate the agent instructions that drive behaviour
- Configure the agent loop: model selection, turn limits, reasoning effort, tool permissions
- Debug agent behaviour in production — wrong tool, wrong arguments, no tool call, early stop - Deploy with Docker to cloud runtimes and instrument runs so they can be debugged after the fact
- Add new features and integrations to live agents without breaking what's already running
- Explain how an agent works to internal teams and, occasionally, to a client's engineers
About you
- 3+ years building production backend software, with strong TypeScript and Node.js
- Hands-on experience building LLM agents that call tools — any framework (Claude Agent SDK, OpenAI, LangChain/LangGraph, Vercel AI SDK, or your own loop)
- Practical experience with MCP or equivalent tool-integration patterns
- Solid REST API integration experience against third-party systems
- Comfortable with Docker and at least one cloud platform (AWS, GCP or Azure)
- Git, testing and code review as normal working habits
- Makes decisions and takes initiative. You choose the model, the flow and the tool surface without being told, flag problems nobody has noticed yet, and propose fixes — including to infrastructure you don't own
- Clear written communication in English
Nice to have
- Claude Agent SDK or the Anthropic API in production
- Writing MCP servers, not just consuming them
- Evaluating LLM output systematically — test sets, regression checks, eval harnesses
- AWS hands-on: ECS/Fargate, Lambda, IAM, DynamoDB
- Customer service platforms — Gladly, Zendesk, Salesforce, Amazon Connect
- Python for data work
