Services practice for the ClickHouse and Langfuse stack

ClickHouse underneath. Langfuse on top. A runbook at the end.

ClickHouse deployed the supported way, on ClickHouse Cloud or inside your VPC. Langfuse live on real production traffic. Analytics workloads moved off the stores that bill by the gigabyte.

  • Fixed price
  • Fixed scope
  • Recorded handover
TRACEA8F3C1E2PROD0 ms · $0.0000
deal_review.run
crm.retrieve_context
agent.strategist
bedrock.claude-sonnet
agent.scorecard
eval.judge
clickhouse.insert
sinkclickhouse.observationsengineReplicatedMergeTreeretentionTTL 90d · ingest lag 400 mscacheprompt cache 0 / 1 hit · caching candidate

ClickHouse

Deployed the supported way

ClickHouse Cloud, or BYOC inside your VPC, from infrastructure as code committed to your repo. Private networking, a tested restore, and retention designed in so storage growth is bounded from day one.

Langfuse

Live on real production traffic

Tracing that survives async and sub-agent boundaries, cost attributed per agent and per workflow, evaluators calibrated against your own experts, and prompts under version control.

Migrations

Off the per-gigabyte meter

Analytics and observability workloads moved from warehouses and log stores onto ClickHouse, starting with a fixed-price assessment that says in writing whether the move is worth it.

Certified on

  • ClickHouseCertified Developer
  • AWSCertified
  • Microsoft AzureCertified
  • SnowflakeCertified
  • Google CloudGenerative AI Leader

Fixed price · fixed scope · recorded handover

Pick a scope. Get a fixed price. Keep the runbook.

Every engagement is a fixed price for a fixed scope, with no hourly meter, so the risk of an overrun sits with us. The Langfuse scopes run on Langfuse Cloud or on an instance you already host. The ClickHouse scopes build, watch and move the database underneath.

Langfuse

Runs on Langfuse Cloud, or on an instance you already host

1 week

Langfuse Launch

Tracing and cost on real production traffic, for one service.

You get

  • Span taxonomy, OTEL propagation across sub-agents
  • Cost per agent and per workflow
  • Custom model definitions for Bedrock and Vertex
  • PII masking client-side, dashboards, cost alerts to Slack or webhook
1 week

Langfuse Evals

Two evaluators, calibrated against labels from your own experts.

You get

  • Two scored dimensions, each with its own rubric
  • Judge calibrated against your experts’ labels
  • Sampling rate set before it runs on all traffic
  • Annotation queue and evaluator dashboards
2 to 3 days

Langfuse Prompts

Prompts out of the repo and under version control, so a wording change is not a deploy.

You get

  • Prompts versioned in Langfuse, out of the repo
  • Labels per environment, rollback without a deploy
  • Prompt-to-trace linking so changes are measurable
  • First experiment run with your team
Half day

Langfuse Workshop

A working session on your own traces, for the engineers who will own the instance.

You get

  • Your traffic, not a demo dataset
  • Taxonomy and evaluator priorities agreed in the room
  • Unit economics: observations, scores, the bill
  • Recorded, with exercises and notes left behind

Bundle

Langfuse Full Loop

Launch, Evals and Prompts, plus the workshop, delivered as one scope. Tracing and cost land in week one, evaluation in week two, and prompt management alongside, so the loop closes before we leave.

ClickHouse

Deploy it, put observability on it, keep it healthy, or decide whether to move onto it

2 to 4 weeks

Langfuse on ClickHouse

Self-hosted Langfuse on ClickHouse Cloud, or BYOC inside your VPC, from IaC committed to your repo.

You get

  • ClickHouse Cloud or BYOC provisioned from committed IaC
  • Private networking, backups, a tested restore
  • Retention designed in: Data Retention (Enterprise) or TTL plus S3 lifecycle rules
  • Operations runbooks written for your on-call
3 to 4 weeks

Observability on ClickHouse

Logs, traces, metrics and session replay landing in ClickHouse through OpenTelemetry, with HyperDX on top. Where Datadog and Splunk workloads go.

You get

  • Managed ClickStack on ClickHouse Cloud, or self-hosted
  • OpenTelemetry collector, ClickHouse, HyperDX
  • Schema, sort keys, TTL and rollups sized to your volume
  • Dashboards and alerts rebuilt from the ones you use
2 weeks

ClickHouse Health Check

For a ClickHouse that is live and struggling, under Langfuse or anything else: slow, expensive, or hours behind real time.

You get

  • Ingestion lag, merges, parts and storage growth
  • Sort keys, partitions and TTL reviewed against real queries
  • Top queries benchmarked before and after, numbers published
  • Prioritised written findings your team can act on
2 to 4 weeks

Migration Assessment

Whether moving off your warehouse, log store or self-managed cluster is worth it, in writing.

You get

  • Source inventory with complexity scoring
  • Cost baseline against a modelled ClickHouse target
  • Schema mapping, including what has no equivalent
  • Phased roadmap, validation harness, and a price for the move

After handover

Expert Retainer

Advisory on instrumentation, cost and operations for a stack your team runs. Async questions answered the same business day, a monthly review call, and a second pair of eyes before a model, retention or schema change ships.

Langfuse on ClickHouse

Self-host when one of three forces crosses a threshold.

Self-hosting Langfuse on ClickHouse is the right answer when one of these forces crosses a threshold. Any one is sufficient. Two together make it inevitable.

ClickHouse Cloud
The default. Scale tier is the floor
ClickHouse BYOC
Data plane in your account, for mandated residency
Self-managed
Free to licence, expensive in people
  1. Cost

    Langfuse bills traces, observations and scores as units. A tool-using agent at roughly 20 units per conversation reaches the fork far sooner than a chatbot at 2.

  2. Residency

    Legal or security decides traces contain regulated data that cannot leave the perimeter. Langfuse Cloud runs in the US, the EU, Japan and a HIPAA US region; a requirement for anywhere else lands here.

  3. Joins

    You already run ClickHouse and want traces sitting next to product analytics and billing, queryable in one place.

Only need network isolation? AWS PrivateLink is available on Langfuse Cloud, on the Enterprise plan with a committed contract. Self-hosting is a much bigger commitment than most isolation requirements need, and we would rather say so than sell you a deployment.

ClickHouse Migrations

We do not quote a migration. We quote an assessment.

Seven sources, one destination, and a written go or no-go before anything moves. The assessment mines your query history rather than the ER diagram, models storage and compute at your real volume, and lays out a phased roadmap with a rollback position at every phase. If the answer is stay, that is the report.

Two to four weeks, fixed price. Open-source ClickHouse to ClickHouse Cloud sits at the bottom of that range; Splunk or Datadog to ClickHouse at the top. The spread is source complexity, not discount. Where a feature has no ClickHouse equivalent, Datadog synthetics and RUM for instance, the report says so.

Book a scoping call
Elasticsearch
Splunk
Datadog
Snowflake
BigQuery
Redshift
Self-managed ClickHouse

Fixed-price assessment

Written go or no-go

2 to 4 weeks. Inventory, cost model, schema mapping, roadmap, and a price for the move.

inventory
query mining
schema map
cost model
rollback
verdict

ClickHouse Cloud

or BYOC in your VPC. Log and trace workloads land in ClickStack. Cutover follows a shadow run.

illustrative run · not client data

What you get at the end

  • Workload assessment report with a full source inventory and per-object complexity scoring
  • Current-state cost baseline and a modelled ClickHouse target-state cost, compute and storage separated
  • Concept and schema mapping, including the features with no ClickHouse equivalent
  • Phased roadmap with an effort estimate and a rollback position at every phase
  • Validation harness specification with the parity checks agreed in advance
  • A fixed price for the migration itself, phase by phase

How an engagement runs

Size, build, instrument, hand over.

Four stages, in this order, every time. The order is the point: nothing is provisioned before it is sized, and nothing is handed over without a runbook.

  1. 01

    Size before you build

    We verify access and prerequisites, size observation volume from the shape of your real traffic, and choose the topology before anything is provisioned. Trace depth, not user count, decides the bill.

  2. 02

    IaC-first infrastructure

    Every cluster, network rule and Langfuse deployment lands in version-controlled infrastructure as code in your repository. You can rebuild it without us, which is the point.

  3. 03

    Instrument what matters

    SDK or OTEL wiring into your real services: spans, sessions, cost tracking with custom model definitions so Bedrock and Vertex numbers are correct, client-side PII masking, and propagation across sub-agents.

  4. 04

    Hand over, don’t hand off

    Every engagement closes with a recorded 90-minute session with your on-call in the room, plus written runbooks for retention, partition reclamation and ingestion-lag triage. Nothing sits in a Bonneville account.

Case study · Ruby AI · Langfuse Cloud activation

Forty-plus agents after every sales call. Every one of them traced.

Ruby, an AI sales-coaching platform, fires more than forty agents in parallel after every sales interaction. Before Langfuse, nobody could trace what each agent did, why outputs varied, or where the spend went.

We built Ruby's observability layer the same way we build it for clients, and we still carry its pager. Cost is attributed by workflow and by customer segment, model changes are evaluated against evidence, and the dashboards stay responsive as volume grows.

Platform
Multi-agent sales coaching
Deployment
Langfuse Cloud
Instrumentation
SDK + OTEL, propagated across sub-agents
Cost attribution
Per workflow, per customer segment
Status
In production, pager held by us
Ruby AI · multi-agent LLM platform
  • Span coverage across async sub-agents
    0%100%
    +100 pts
  • Agents traced per sales interaction
    040+
    all of them
  • Where LLM spend shows up
    invoiceper trace
    every dollar owned

Cost attributed per workflow and per segment. Dashboards and cost alerts live from the first week of production.

Questions we get on the first call

Straight answers, including the ones that cost us the deal.

Langfuse Activation

  • Stay on it. Langfuse Cloud is the right answer for most teams and we do not compete with it. Signing up is easy; what stalls is everything after tracing: nobody can say whether last week’s prompt change made answers better or worse, which evaluator to build first, or what a single agent run costs, and the prompts are still in the repo so every wording change is a deploy. Every Langfuse Activation scope runs on Langfuse Cloud.

  • Yes. Langfuse’s documentation is good, their workshop is public, and nothing we do is secret. What you are buying is the week. Someone who has already made the taxonomy mistakes, already calibrated a judge against human labels, and already reconciled cost figures to a provider invoice does in five days what a first-timer does in six weeks of nights and weekends.

  • A billable unit is one trace plus every observation plus every score, counted on ingestion. Two consequences people miss. Units Langfuse generates itself count, including LLM-as-a-judge output, annotation labels and experiment runs. And because units are counted on ingestion, deleting data never refunds a unit: retention reduces storage cost, not unit cost. We project unit volume before evaluators switch on across all traffic.

  • A recorded 90-minute session, a written runbook, and every artifact left in your project and your repo: the span taxonomy, the dashboards, the evaluator rubrics, the judge sampling rate and what it implies for your bill, and the alert configuration. Nothing sits in a Bonneville account. The test we hold ourselves to is whether it survives the engineer who leaves.

  • No. On self-hosted Langfuse, SSO and SSO enforcement are free in the open source build. This is the most common misconception in the stack and it costs teams real money. The Enterprise licence buys governance: project-level RBAC, audit logs, the admin and instance management APIs, data retention management, server-side masking, and the SOC 2 and ISO 27001 reports procurement will ask for. On Langfuse Cloud it works differently: SSO requires the Teams add-on on Pro, or the Enterprise plan.

ClickHouse

  • You can, and it will fail quietly. Everything loads, the counts match, and the queries are slower than the source. The method that works is to mine the query history rather than the ER diagram, derive the schema from what people actually run, go wide on one event table per business process, and push dimension attributes onto the event at write time. The sort key is the expensive part to get wrong: ORDER BY and PARTITION BY are one-way doors, and changing them later means rebuilding the table.

  • Almost always one of three things, and a Health Check tells you which. The cluster was sized from trace count rather than observation count, so storage lands at ten or a hundred times the model. Retention never reclaims: lightweight deletes mark rows rather than removing them, and monthly partitions above the merge threshold are never rewritten, so deleted data keeps occupying disk; one publicly documented deployment carried 679 GiB of unreclaimed deleted rows in a single monthly partition. Or there is no TTL baseline and no storage tiering, so everything stays on hot storage forever.

  • No. We quote an assessment, and the migration is quoted from its findings. Most partners sell you a migration; this engagement tells you whether to do one. If it concludes that two of your five workloads should stay where they are, that is in the report. Assessment and code conversion are the hard parts of a migration and moving the data is the easy part, so skipping assessment is how teams discover the schema was wrong three months in.

  • Not always, and we will not promise it before seeing 90 days of your account usage. ClickHouse wins decisively on storage cost against Elasticsearch and Splunk, where you pay for data you rarely query, and on query cost under high concurrency. It can be more expensive than Snowflake on an idle-heavy workload that runs twice a day, because auto-suspend is genuinely hard to beat there. Scale-to-zero on idle services is the closest analogue, and whether it closes the gap depends on your tolerance for wake latency.

  • No. We run a shadow. Your existing system stays the system of record, the same data lands in ClickHouse in parallel, and the queries run against both on a schedule with automatic diffing. You get results-match-or-not, p95 on both sides, monthly cost on both sides, and an engineering-hours number. Nothing customer-facing changes, so during the shadow there is nothing to roll back. Cutover is a separate decision made with all of that in hand.

About Bonneville Data

Thirty minutes to a fixed-price scope.

A delivery practice for the ClickHouse and Langfuse stack, working remotely across North America. We deploy and instrument the stack, migrate teams onto ClickHouse, and run the same stack in production ourselves rather than on client accounts only.

Start with a 30-minute scoping call. You leave it with the scope that fits and a named start date.

What happens next

Scoping call
30 minutes, with an engineer
Written scope and price
Within two business days
Start date
Named on the call
Bring
Trace volume, sources, data residency