Scroll

Underlight
AI Systems

We build AI systems that survive contact with production — agents that finish the task, retrieval that cites its sources, and evaluation that tells you the truth. We work beneath the surface, so the product on top feels effortless.

[001]

About

We're an applied AI studio. We build AI systems for companies that want to move faster, automate manual work, and scale without losing accuracy.

The hard part is rarely the model. It's retrieval quality, evaluation, guardrails, cost per call, and deciding who reviews an answer the system isn't confident about. That is the part we build.

[002]

Capabilities

Model-agnostic by design. The retrieval layer, the orchestration, and the evals are what compound — swapping a model should be a configuration change, not a rewrite.

01

Models & Reasoning

Anthropic Claude
OpenAI
Google Gemini
Llama
Mistral
Hugging Face
Ollama
02

Agents & Orchestration

Model Context Protocol
LangChain
Temporal
Ray
Airflow
n8n
03

Data & Retrieval

PostgreSQL / pgvector
Qdrant
Elasticsearch
Kafka
Snowflake
Databricks
Redis
04

Platform & Delivery

Python
PyTorch
TypeScript
FastAPI
Docker
Kubernetes
NVIDIA
AWS
Terraform
Weights & Biases
Grafana
Vercel
[003]

Work

The kind of systems we build, and what it actually takes to put them in production.

Insurance · Agent systems
2026

Claims Triage

An agent pipeline that reads a claim, extracts the facts, and routes it — with a reviewer on top.

Legal · Retrieval
2026

Contract Intelligence

Retrieval over a contract estate that answers with clauses, not vibes.

Manufacturing · Computer vision
2026

Shop-Floor Vision

Defect detection at line speed, on hardware that already exists in the plant.

SaaS · Assistants
2026

Support Copilot

An assistant inside the support console that drafts, cites, and knows when to stop.

Process
[004]

Process

Four phases, each with something working at the end of it. No discovery deck that leads to another discovery deck.

011–2 weeks

Discover

We sit with the people who do the job today and write down how it actually works, including the exceptions nobody documented. Then we build the evaluation set and agree what good has to score before anything ships.

Workflow mappingData auditEval designFeasibility
01
023–6 weeks

Prototype

A working system in front of real users inside a month. We run against the eval set every day, keep a scoreboard everyone can see, and kill approaches early when the numbers say to.

Retrieval & agentsEval harnessWeekly demosCost modelling
02
034–8 weeks

Deploy

Production means guardrails: scoped permissions, human review where it matters, fallbacks when a provider degrades, cost ceilings, and tracing on every call. We ship behind a flag, to a slice of traffic, and watch it.

GuardrailsObservabilityLoad & cost testingStaged rollout
03
04Ongoing or handover

Operate

Models change, data drifts, and the business moves. The system ships with the monitoring and the evals needed to upgrade a model safely, long after the build is finished.

MonitoringModel upgradesRetrainingTeam handover
04
[005]

Contact

Tell us about the workflow, the data, and what breaks when the answer is wrong. We'll tell you whether AI is the right tool.

© 2026 UnderlightAll rights reserved.Remote-first · Europe & MENA