About the role
Not a prompting job. Our agent does real work in production: holds sessions for weeks, runs scheduled jobs, executes code, acts through Gmail, Calendar, Notion, Slack and GitHub. You own the whole loop: tools, workflows, memory, routing across 50+ models with mid-stream fallbacks, and the evals that prove it works. No research layer between you and users.
What you'll do
- Ship agent tools and the workflows that chain them into real outcomes.
- Own routing, fallbacks and latency across 50+ models.
- Build evals that catch regressions before users do.
- Take new models from release post to production the same day.
What we're looking for
- You've run agents or tool-calling systems in production at scale, and can walk through exactly what broke and how you caught it.
- You think in failure modes: provider outages, rate limits, silent regressions, corrupted memory.
- You benchmark before you believe anything, including your own work.
- Strong TypeScript. You pick up anything else in days, not weeks.
Nice to have
- MCP or similar connector ecosystems in production.
- Inference optimization or model serving.
Apply for this role
AI Engineer · Remote
More in Engineering