Skip to content

How we work

How we work, step by step.

Five stages from first conversation to a system your team runs without us. Each one has a clear output, and a point where you can stop if it isn't working.

01

Discovery

1–2 weeks

We sit with the team that will use the system and the team that owns the data feeding it. The goal is one sentence: what decision does this support, and what number tells us if it's working.

  • Interview stakeholders from the business side and the data team separately, then reconcile what each expects the system to do.
  • Pull sample data extracts and check for the failure modes that kill projects later: survivorship bias, restated financials, missing point-in-time snapshots.
  • Define the metric that decides success, and the threshold below which we recommend stopping.
  • Write the one-page scope and get sign-off from the budget owner and compliance before architecture work starts.

What you get: A written scope: the decision, the metric, the data sources, and what is explicitly out of bounds for phase one.

02

Architecture

2–4 weeks

We design the data layer before the model: where data comes from, how point-in-time correctness is enforced, what the model can and can't see, and where a human has to sign off.

  • Map every data source to its owner, refresh cadence, and licensing constraints.
  • Design the point-in-time data layer so a backtest can never see information that wasn't available on the date in question.
  • Decide where a human approval gate sits in the workflow, and what triggers it.
  • Run a go/no-go architecture review with your engineering and compliance leads before a line of production code is written.

What you get: An architecture document and a go/no-go review with your engineering and compliance stakeholders, before any production code is written.

03

Pilot

4–8 weeks

We build against real data with evaluation running from day one, not bolted on at the end. The pilot has to beat the current process on the metric agreed in discovery, or we say so plainly.

  • Build the first working version against real data, not a synthetic sample.
  • Stand up the evaluation harness before the model, so every version gets scored the same way from day one.
  • Run the pilot's output against the current process side by side, on the metric fixed in discovery.
  • Deliver an honest recommendation, proceed, iterate, or stop, backed by the evaluation report.

What you get: A working system on real data, an evaluation report, and a clear recommendation on whether to proceed.

04

Production

6–12 weeks, engagement-dependent

We harden the pilot: monitoring for drift and cost, an audit trail, sign-off steps where mistakes are expensive, and integration into the systems your analysts already use.

  • Add monitoring for model drift, data quality breaks, and per-query cost, with rollback rehearsed before go-live.
  • Build the audit trail: every output traceable back to its inputs and the model version that produced it.
  • Wire sign-off steps into the workflow wherever a mistake would be expensive.
  • Integrate into the tools your analysts already have open, not a new tab they have to remember to check.

What you get: A production deployment inside your environment, on-prem or VPC, with monitoring and rollback in place before go-live.

05

Handover & enablement

2–3 weeks

We document the system, hand over the evaluation suite and runbooks, and spend time with the people who'll own it, so it doesn't quietly become tribal knowledge only we have.

  • Document the system: architecture, data lineage, and the reasoning behind each guardrail.
  • Hand over the evaluation suite so your team can rerun it after every model or data change.
  • Run working sessions with the people who'll own the system, rather than a farewell demo.
  • Agree a defined support window, after which the dependency on us formally ends.

What you get: Documentation, an eval suite your team can rerun, and a defined support window before we step back.

Engagement models

Four ways to bring us in.

Each is scoped to a clear outcome and rolls into the next only when that makes sense, not before.

Forward-Deployed Team

Embedded AI engineers inside your workflows, shipping production-grade tools your finance team actually uses. We work from inside your sprint cadence, not alongside it.

Project Delivery: Time & Materials or Fixed-Scope

Time & Materials, when scope evolves as you learn. It suits exploratory model builds and roadmap-defining pilots. Fixed-Scope, when requirements, architecture and compliance constraints are already defined and the deliverable is known upfront.

AI Advisory Retainer

Ongoing strategic guidance on model risk, vendor selection, and AI governance: a standing line to a senior engineer for the decisions that come up between projects.

Build-Operate-Transfer

We stand up the team and capability, then hand over a fully operational AI function to your own staff, on a timeline agreed at the start rather than an open-ended dependency.

Collaboration principles

What stays true across every engagement.

You talk to the builders

You work with the people writing the code, not an account manager relaying it back and forth.

Show the working

Every output carries its sources. In finance, an answer you can't trace is an answer you can't use.

Stop points, not lock-in

Each stage ends with a clear go/no-go. Nothing rolls into the next phase automatically.

We leave you able to run it

Handover isn't a slide. It's documentation, an eval suite, and time with the team who'll own the system after we're gone.

Want to see how this maps to your stack?

Book a 30-min intro call