Four things we are asked for, over and over
Small language models
A compact model trained on your own documents, tickets and product
data. It runs on hardware you control, costs a fraction of a frontier
API call, and does not leak anything to a vendor. We handle data
preparation, fine-tuning, evaluation and the serving stack.
Fine-tuning, distillation, evaluation harnesses
Agents that finish the job
Most agent demos fall apart on the second turn. We design agents
around real tools and real permissions, with the retries, guardrails
and audit trail that let a business actually depend on the output.
Tool use, retrieval, guardrails, evaluation
Orchestration you can operate
One agent is a prototype. Twelve agents sharing state, budget and
failure modes is a system. We build the routing, memory and
observability layer so your team can debug it at 2am without calling
us.
Multi-agent routing, state, tracing, cost control
The data underneath
Features, pipelines and warehouses that feed the models. Usually the
least glamorous half of the work and the half that decides whether
anything above it survives contact with production.
Feature stores, pipelines, warehouse modelling
How an engagement runs
-
Diagnostic
Two weeks. We read the data, talk to the people using it, and come
back with what is worth building and what is not. You keep the
write-up either way.
-
Prototype
Four to six weeks to a working thing on your data, measured against
a metric you agreed to before we started.
-
Production
Serving, monitoring, cost controls, rollback. The unglamorous
months where a demo becomes something on call.
-
Handover
Your engineers own it. We document, pair, and step back. A good
engagement ends with you not needing us.
Twenty years of shipping data systems, most of them
before anyone called it AI.
BrainField is led by an engineer who has spent two decades building
data and machine learning platforms for large retail and e-commerce
organisations — feature stores serving live traffic, warehouse
migrations, recommendation systems, and multi-agent analytics
platforms on Google Cloud and AWS.
That history is why we tend to argue for the smaller model, the
boring pipeline and the shorter dependency list. It is usually what
is still running two years later.
What we work in
- Google Cloud, Vertex AI
- AWS, SageMaker
- BigQuery, Snowflake
- Vertex Feature Store
- PyTorch, Hugging Face
- LangGraph, MCP
- Airflow, dbt
- Looker, LookML
- Kubernetes, Cloud Run
- Python, SQL, Java