AI Engineer

  • Permanent
  • Full time
  • Remote

Job Purpose and Summary

This position is responsible for designing, building, and maintaining the AI systems: the LangGraph agents, retrieval-augmented generation (RAG), multi-provider LLM orchestration, and prompt/evaluation machinery that turn customer conversations and uploaded documents into grounded, auditable cost estimates. Good leadership skills and ability to work in US time zones as required, are both pre-requisites.


Focus Area(s) for Impact:

Retrieval grounding & estimate/sizing accuracy
Agent reliability, latency, and LLM cost-efficiency

Why choose us?

Engineering at QED is driven by curiosity, responsibility, and real constraints. We work on real-world problems that demand clear thinking, strong engineering judgement, and ownership from start to finish. We focus on building systems that need to work in production and stand the test of time not on buzzwords or shortcuts.

Our projects span demanding international environments and are built using disciplined engineering practices. We invest time in understanding problems properly and designing with intent, and delivering solutions that hold up under real usage.

We've built an environment that promotes collaboration, trust, quality and strong systems rather than individual heroics.

Clear, factual communication matters to us because clarity is kind. We share thoughts openly, prioritising facts for learning. We stay curious and challenge assumptions every day to continuously raise our standards. Brilliantly Human. Frontier-Focused.

Everyone has real influence over architecture, technical direction and standards. We make decisions close to the work, balancing autonomy with accountability. Exceptional talent grows alongside experienced engineers through real problems.

We support flexible ways of working, long term development and prioritise sustainability. If you’re motivated by high standards and meaningful ownership, we believe QED is a place where you can do your best work.


Key Accountabilities & Responsibilities


LLM Agents & Orchestration
Build and maintain LangGraph agents: ReAct retrieval sub-agents, custom middleware (call-limit, completion signaling), Pydantic-typed graph state, and streaming flows.
Metrics / Measure of Success: Agent task success rate.

Retrieval-Augmented Generation
Own retrieval quality end to end: agentic retrieval with collection routing, LLM-driven metadata filters and similarity thresholds, query rewriting, semantic chunking, and citation tracking.
Metrics / Measure of Success: Retrieval precision/recall on the eval set; grounding/citation accuracy. Groundedness score on the standard eval set.

LLM Provider Integration
Providers behind the abstract interface: AWS Bedrock, OpenAI, self-hosted vLLM.

Backend Services & Data
The platform's API services, relational and domain-knowledge data stores, the vector store, and observability/tracing for AI workflows.

Technical Leadership & Collaboration
Set technical direction for the AI workstream; review designs and code, mentor engineers, and align with product stakeholders. Break the status quo.
Metrics / Measure of Success: Delivery against roadmap; quality of reviews and design decisions; growth and unblock rate of mentored engineers; stakeholder satisfaction.

Education and Experience

Formal Education

Required: Bachelor's degree in Computer Science, Software Engineering, Data Science, or Mathematics.

Desirable: Master's degree in Computer Science, Machine Learning, or AI; relevant cloud or machine-learning certification.

Experience

Required: Minimum of 6 years of progressive software engineering experience, including 1+ years building production LLM-powered or machine-learning systems (agents, RAG, prompt/evaluation pipelines). Proven delivery in a fast-paced environment and availability to overlap U.S./Eastern-Time business hours for client meetings.

Desirable: Track record of shipping AI features to production quickly with measurable quality, latency, or cost improvements, and of working directly with clients.

Required Competencies and Skills

Technical Proficiency: Expert in production software engineering and in building LLM-powered systems: agents/orchestration, retrieval-augmented generation, and prompt/evaluation pipelines; fluent in using AI development tools to accelerate delivery.

Functional Expertise: Deep understanding of applied AI engineering: retrieval and grounding, embeddings and vector search, multi-model integration with cost/latency tuning, and evaluation-driven iteration. Working knowledge of the cost-estimation / Work Breakdown Structure domain, or the ability to learn a deep domain quickly.

Problem-Solving: Ability to analyse complex, non-deterministic AI behaviour from real production traces and deliver scalable, well-tested solutions fast under time pressure (e.g., diagnosing over-retrieval, hallucination, or estimate/sizing collapse and landing a measurable fix).

Communication: Exceptional ability to communicate technical and model-behaviour concepts clearly to non-technical and client stakeholders; confident in live U.S./ET client meetings and disciplined in clear, async written updates across time zones.

Judgment: Proven ability to make critical, time-sensitive decisions with limited information: balancing accuracy, latency, cost, and risk; high ownership and motivation, with a consistent bias to action and willingness to go the extra mile to deliver.

Optional / Nice to Have Skills

Technical Proficiency: Self-hosted model inference; observability and tracing tooling for AI workflows; clean, layered service architecture; building internal AI tooling/automation that speeds up the whole team.

Functional Expertise: Cloud-platform depth; agent-reliability patterns (call caps, guardrails, convergence control); ingestion and processing of documents across multiple formats.

Problem-Solving: Experience designing evaluation/benchmark frameworks and regression fixtures that hold AI quality steady over time.

Communication: Experience working directly with U.S.-based clients; mentoring engineers, leading design reviews, and representing AI work to clients and leadership.

Judgment: Experience setting technical direction and making roadmap trade-offs for an AI workstream or team; thrives in a high-growth, fast-delivery environment.

|