All rolesSenior TypeScript Engineer — LLM Agents
Full remote · Full-time · Independent contractor
We are looking for a senior backend engineer to design and deliver production AI-agent and retrieval systems. The role covers the complete engineering lifecycle: clarifying client workflows, defining system boundaries, implementing reliable orchestration and integrations, establishing evaluation and observability, and improving behaviour after release. This is production software engineering rather than model research or prompt-only prototyping. You will work directly with delivery specialists, engineers, and clients to turn ambiguous processes into reliable, measurable, and maintainable products.
What you will do
- Design and deliver production agents and multi-step LLM workflows in TypeScript and Node.js.
- Define reliable tool contracts, structured outputs, retries, timeouts, approval boundaries, and human escalation.
- Build retrieval pipelines covering ingestion, chunking, embeddings, indexing, retrieval, reranking, grounding, and citations.
- Develop backend APIs and services with NestJS and PostgreSQL.
- Create evaluation datasets and automated checks for correctness, tool use, retrieval quality, latency, and cost.
- Instrument traces, logs, metrics, token usage, and model costs.
- Investigate production failures and improve prompts, retrieval, tools, and orchestration using evidence.
- Integrate model providers behind stable application interfaces and support provider or model changes.
- Participate in client discovery, explain trade-offs, estimate work, and deliver in reviewable increments.
- Review human- and AI-generated code to the same engineering standard.
Required experience
- 5+ years building and operating production backend or full-stack software.
- Strong TypeScript and Node.js, including API design, asynchronous programming, testing, and production debugging.
- At least one substantial LLM, agent, tool-calling, or production RAG system delivered beyond the prototype stage.
- Practical experience with tool/function calling and either agent orchestration or retrieval.
- Strong PostgreSQL, SQL, and data-modelling skills.
- Ability to explain real failure modes including hallucination, poor grounding, brittle tools, loops, latency, cost, and security.
- Production ownership including CI/CD, monitoring, incident investigation, and incremental delivery.
- Strong written and spoken English for client and team communication.
Useful, not required
- NestJS.
- LangGraph, Vercel AI SDK, Mastra, OpenAI Agents SDK, or custom orchestration.
- LangSmith, Langfuse, OpenTelemetry, or equivalent evaluation and tracing tools.
- Qdrant, pgvector, Pinecone, Weaviate, or another vector store.
- MCP client or server integrations.
- Python for experiments and data-processing tasks.
- Previous consultancy, outsourcing, or direct client-facing delivery.
What we offer
- Full-remote work with a flexible schedule outside agreed client and team overlap.
- Full-time independent-contractor engagement.
- 14 paid vacation days and 7 paid sick days per year.
- Official public holidays in your country of residence in addition to the paid days above.
- Work from your own laptop; EthBerry provides the project accounts, model access, and paid software licences required for the engagement.
- Compensation agreed individually based on relevant experience.
- Direct participation in architecture and technical decisions on client projects.
- The contract does not include health insurance, company hardware, or office-based benefits.
Hiring process
- Application review.
- HR screening — 30 minutes.
- Technical interview — 75 minutes. We discuss one production system, its architecture, failures, evaluation strategy, retrieval quality, and operational trade-offs.
- Unpaid take-home assignment — strictly capped at 4–6 hours. Build a small TypeScript agent with two tools, retrieval over supplied fixtures, one malformed tool response, one no-answer case, a compact evaluation set, and structured logs. No deployment, frontend, personal API spending, or production-ready polish is required.
- Assignment review — 75 minutes. You explain decisions and respond to one changed requirement.
- Final cultural-fit interview — 30 minutes.
AI coding tools are allowed. Their use must be disclosed briefly, and you must be able to explain and defend every submitted decision. The assignment is not used in production.
Apply