AI features that ship with a cost ceiling, a human decision path and an audit trail

For product teams adding assistants, document intelligence, grading or automation to software that already has users. We design the cost ceiling, the fallbacks and the human review step before the prompts, so the feature stays affordable at scale and defensible to a regulator.

Talk to us if

  • Users are asking for an assistant and finance is asking what it will cost.
  • A prototype works in a demo and nobody will sign off on putting it in front of customers.
  • The model is confidently wrong often enough that support is paying for it.
  • A regulator or a client will ask why a decision was made, and you have no answer.
  • Your AI bill grows with usage while your subscription revenue does not.

Who this is for

  • A product with users asking for an assistant and a finance team asking the cost
  • A hiring, admissions or grading workflow that must keep a human as the decider
  • A team with a working prototype that needs to be production-safe
  • A business sitting on documents nobody has time to read

What we build

  • Assistants and copilots

    Answers grounded in your own content, with citations, so a user can check the source instead of trusting the tone.

    • Retrieval over your documents, tickets and product data
    • Citations back to the source passage on every answer
    • Scoped permissions, so an assistant can never answer from data the user cannot see
    • Refusal and escalation paths when the answer is not in the corpus
  • Document and content intelligence

    Turning documents into structured data your systems can actually use.

    • Extraction into typed, validated schemas rather than free text
    • Classification, routing and summarisation at volume
    • Handling of scans and mixed-quality inputs, with confidence thresholds
    • A review queue for anything below the threshold
  • Grading, screening and review workflows

    Where the output affects a person, the design question is not accuracy but accountability.

    • One human-decider step by construction; the model drafts, a person decides
    • Rubrics and structured scoring instead of an opaque verdict
    • Audit logs recording what the model saw, what it proposed and who decided
    • Bias and adverse-impact checks built into the evaluation set
  • Agents and automation

    Multi-step work that touches real systems, built with the failure modes designed in.

    • Tool use against your own APIs, with typed arguments and validation
    • Queues, retries and idempotency, so a retried step does not double-charge anyone
    • Step limits, timeouts and a kill switch
    • Human approval gates on anything irreversible
  • Cost control and reliability

    The part most teams add after the bill arrives. We build it first.

    • A published cost model per journey before a prompt is written
    • Hard allowances per tenant and per user, with a clear message when they run out
    • Token budgets per request, prompt caching and model routing by task
    • Circuit breakers and a non-AI fallback path when a provider degrades
  • Evaluation, privacy and safety

    How you know a prompt change made things better rather than different.

    • Evaluation sets that run in continuous integration on every change
    • Regression tests on the cases that previously went wrong
    • Prompt-injection and jailbreak testing against your own surfaces
    • Data minimisation, redaction and a documented retention position

What you get

  • A cost model and allowance design before any prompt is written
  • The feature itself, with structured outputs and validation
  • An evaluation set wired into CI so quality is measured, not felt
  • Rate limits, circuit breakers and a degradation path
  • Audit logs and a written note on data handling for your legal review

What you gain

  • A feature whose cost per user you can state before you launch it
  • Quality tracked as a number that moves, not a feeling
  • An answer for the regulator, the client or the board about how a decision was made
  • A graceful degradation instead of an outage when a provider has a bad day

How the engagement runs

  1. Cost and risk assessment

    What it costs at ten times current usage, and what happens to the user when the budget runs out.

  2. Prototype with an evaluation set

    A narrow version measured against real cases, so quality is a number rather than an opinion.

  3. Production hardening

    Allowances, fallbacks, audit logs, permissions and the human decision step.

  4. Run and measure

    Cost per journey, quality trend and failure rate, reported monthly.

Technology we work with

Models and providers
  • Claude API
  • OpenAI API
  • Amazon Bedrock
  • Google Vertex AI
  • Open-weight models
Retrieval
  • pgvector
  • PostgreSQL full-text search
  • Embedding pipelines
  • Hybrid search and re-ranking
Orchestration
  • TypeScript
  • NestJS
  • BullMQ
  • Structured outputs and tool use
  • Streaming responses
Evaluation and observability
  • Evaluation harnesses in CI
  • Token and cost metering
  • OpenTelemetry
  • Prompt and response logging
Screens from Jversity.
Education · 2026–ongoing

Building Jversity: a multi-tenant educational LMS platform

Live classes, courses, interview preparation, blogs, video recordings and much more built and operated by us as our own product.

Questions we get asked

Which model should we use?

Whichever the task and the budget justify. We default to the Claude API for text and structured output, and we build the integration so a model can be swapped without rewriting the feature.

How do you stop AI cost scaling faster than revenue?

Hard allowances per tenant and per user, a token budget per request, prompt caching, routing cheap tasks to cheaper models, and a circuit breaker that degrades to a non-AI path. The cost model is published before the build.

Can the AI make decisions about people?

No. For hiring, admissions or grading we design one human-decider step by construction. The model drafts, a person decides, and the audit log shows both.

Will our data be used to train someone else's model?

Not on the configurations we deploy. We use providers and settings that exclude training on your data, and we write that position down for your legal review.

What does success look like for your project?

Tell us what you are building and where it is stuck. We reply within one business day with a straight answer on fit, approach and cost.