Hiring BarSupport

Build vs Buy Framework — Cross-Cutting Pattern

Pattern38 min read6 diagrams

Technologies that implement this pattern: Kafka · DynamoDB · Cassandra · Elasticsearch · Redis · API Gateways · ZooKeeper & etcd

Why This Matters#

Build vs buy is not a technology question. It is a question about where your company is willing to carry a pager for the next five years. Every "we'll just build it" commits a team to an on-call rotation, a security patch cadence, an upgrade treadmill, and a hiring profile — long after the engineer who wrote the first version has moved on. Every "we'll just buy it" commits the company to a vendor's roadmap, pricing model, outage history, and exit cost. Neither is free. The only question is who pays, when, and in what currency.

Most candidates treat build vs buy as a feature comparison: "Kafka has what we need, so self-host it" or "DynamoDB scales, so use it." Staff engineers treat it as a total-cost-of-ownership and differentiation question: does owning this capability make the product measurably better than competitors, and is the 3-year fully loaded cost of owning it lower than the 3-year cost of renting it plus the cost of leaving? Principal engineers go one step further — they treat it as a portfolio decision: the org can afford to be excellent at maybe 3–5 infrastructure capabilities, and every "build" spends from that budget.

In interviews, build vs buy shows up constantly and rarely by name. "Would you use Kafka or SQS here?" "Why not just use Elasticsearch?" "Would you build your own rate limiter?" Each is a build-vs-buy probe wearing a technology costume. The candidate who answers with a feature table gets L5. The candidate who answers with "who runs it, what it costs at 10×, and how we'd get out" gets L6. The candidate who answers with "here's the paved-road decision for the 40 teams that will face this same question" gets L7.

The 60-Second Version#

  • Default to buy for anything that is not your product's differentiator. Auth, email delivery, payments rails, observability, managed databases. If your customers can't tell whether you built it, you shouldn't have.
  • Price engineers, not licenses. A fully loaded engineer costs ~$250–400K/yr. A sustainable 24/7 on-call rotation needs 6–8 engineers. Self-hosting anything with a pager is a ~$1.5–3M/yr commitment before you write a line of feature code.
  • Open source is "free like a puppy." Self-hosted Kafka, Cassandra, or Elasticsearch typically consumes 0.5–2 FTE per cluster family in upgrades, tuning, capacity planning, and incident response. The license is the cheapest line item.
  • The crossover point is real but late. Managed services usually carry a 2–4× markup over raw compute. Owning becomes cheaper only once the managed bill exceeds roughly $1–2M/yr and you have the team to run it. Below that, the markup is the cheapest insurance you'll ever buy.
  • Price the exit on day one. Lock-in is not binary; it is an exit cost measured in engineer-months. A proprietary API with 200 call sites and 50 TB of data can cost 6–18 engineer-months to leave. Write that number down before you sign.
  • Most decisions are two-way doors — data gravity makes them one-way. Swapping an email vendor takes a sprint. Migrating 500 TB off a proprietary datastore takes a year. Reversibility is determined by where the data lives and how many call sites touch the API.

The Problem#

A team needs a capability — a queue, a search index, a feature-flag system, an identity provider, a workflow engine. Three options exist: build it in-house, adopt an open-source project and run it yourself, or rent it as a managed service / SaaS. The feature matrices look similar. The real differences are invisible in a demo: who gets paged when it breaks at 3am, how the bill scales at 10× traffic, what happens when the vendor changes its license or pricing, and how many engineer-months it costs to leave. Teams that decide on the visible 20% (features, list price) routinely pay for the invisible 80% (operations, migration, opportunity cost) for years.


Case Studies That Use This Pattern#

  • Message Queues — Self-hosted Kafka vs MSK/Confluent vs SQS is the canonical build-vs-buy fight
  • Metrics & Monitoring — Datadog bills vs a self-run Prometheus/Mimir stack; the crossover is often the first seven-figure infra decision a company faces
  • Search Indexing — Elasticsearch self-hosted vs managed OpenSearch vs Algolia; operational burden dominates at scale
  • Payment Processing — Almost always buy the rails (Stripe/Adyen), build the ledger; the boundary is the interview
  • Database Selection — Managed Postgres vs self-run vs DynamoDB; lock-in and data gravity drive the choice
  • Rate Limiting — Gateway-provided limits vs a custom limiter platform
  • Distributed Job Scheduler — Build a scheduler vs adopt Temporal/Step Functions/Airflow

The Core Tradeoff#

StrategyWhat WorksWhat BreaksWho Pays
Build in-houseExact fit, full control, no vendor margin, can become a differentiatorScope creep, bus factor, undocumented behavior, 2–3× longer than estimatedThe owning team forever — on-call, upgrades, hiring; product roadmap pays in opportunity cost
Self-host open sourceNo license fee, portable, large community, hireable skillsUpgrade treadmill, tuning expertise, security patching, license changes (Elastic 2021, HashiCorp 2023, Redis 2024)Platform team's on-call; every major-version upgrade is a quarter-long project
Managed open source (MSK, RDS, Cloud SQL, Aiven)Compatible API, provider runs patching/failover, easy exit to self-host2–3× compute markup, lagging versions, limited tuning knobs, shared-fate with provider incidentsFinance pays the markup; app teams inherit provider's maintenance windows
Proprietary managed (DynamoDB, Spanner, Step Functions, Firestore)Near-zero ops, elastic scale, strong SLAs (99.99%+)Proprietary API, data gravity, pricing changes, hard exitThe future team that must migrate; negotiating leverage at renewal
SaaS / vendor product (Datadog, Auth0, Algolia, Stripe)Fastest time-to-value, expertise included, compliance certifications bundledPer-seat / per-event pricing that scales super-linearly, data residency, roadmap dependencyFinance at renewal; security for third-party risk review; users when vendor has an outage
Hybrid: buy now, own the seamSpeed today, exit path preserved via an internal interfaceAbstraction tax — lowest-common-denominator API, 10–20% extra build effortThe platform team maintaining the adapter

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveCompares features and benchmarks of the candidatesAsks whether the capability is a differentiator and who will own it on-callAsks how many teams face this same decision and whether it should be a paved-road default
CostCompares list price vs "free" open sourceComputes 3-year TCO including FTEs, on-call, and migrationPrices it as portfolio spend: headcount diverted from product, vendor concentration, negotiating leverage at renewal
Lock-in"We can always migrate later"Estimates exit cost in engineer-months; wraps proprietary APIs behind a seamDecides which lock-ins the company accepts deliberately (e.g., one cloud) and which it hedges (e.g., observability), and writes it down
FailureTrusts the vendor SLADesigns for the vendor being down: degraded mode, timeouts, fallbackMaps correlated vendor risk across the org — "12 services all depend on the same IdP region" — and sets a concentration policy
Ownership"The platform team will run it"Names the owning team, rotation size, and runbook before adoptingRedraws the boundary: which capabilities are centralized, which teams may self-host, the exceptions process
Time horizonThis quarter's launchThis system's life (2–3 years)3–5 years: license drift, vendor viability, the retire/replace decision
Why "Cost" separates levels

The L5 answer — "open source is free, the vendor charges $X" — is competent but counts only the invoice. The Staff answer adds the people: a self-hosted Kafka cluster with a proper 24/7 rotation is ~1.5 FTE of steady-state work plus spikes for upgrades, which at $300K fully loaded is ~$450K/yr before hardware. The Principal answer adds the opportunity cost: those 1.5 FTE are 1.5 FTE not building the product feature the company is actually competing on, and the question becomes whether the board would fund "run Kafka well" as a strategic initiative. If the answer is no, the managed markup is the cheapest line in the budget.

Why "Lock-in" separates levels

"We can migrate later" is technically true of every system and practically false of most. Staff engineers quantify: 180 call sites × ~2 hours each + a dual-write period + 40 TB of backfill ≈ 6 engineer-months. That number goes in the design doc. Principal engineers recognize that some lock-in is the price of leverage — committing fully to one cloud provider buys enterprise discounts of 20–40% and a single operational model — and they make the accept-vs-hedge call explicitly per capability class rather than letting 40 teams each make it differently.

Why "Ownership" separates levels

A build or self-host decision without a named owner is an orphan waiting to happen. Staff engineers refuse to adopt a system until someone has signed up for the pager, the upgrade plan, and the capacity reviews. Principal engineers look at the org chart and ask whether the right owner even exists: if three product teams each self-host Elasticsearch, the right move may be to fund one platform team and deprecate the other two clusters — a people decision, not a technology one.


Staff Default Position#

Buy what's commodity, build what's differentiating, and own the seam either way.

If the capability is not something customers would notice or pay for, rent it — the managed markup is almost always cheaper than the engineers you'd need. If it is core to your competitive advantage (Netflix's video delivery, Stripe's payment ledger, Uber's dispatch), build it and staff it like a product. In both cases, put a thin internal interface between your business logic and the vendor/system, so the decision stays a two-way door. The default time horizon for the comparison is 3 years, and the default unit is fully loaded cost including on-call, never list price.


When to Deviate#

  • The managed bill crosses ~$1–2M/yr and the workload is stable. At that scale the 2–4× markup buys 3–6 dedicated engineers. Predictable, flat workloads (storage, steady-state compute) are the best repatriation candidates; spiky workloads still favor elastic managed services.
  • The capability is the product or its moat. Latency-critical paths where a vendor's p99 is your p99, or where proprietary tuning is the differentiator (Netflix's Open Connect CDN, Dropbox's Magic Pocket storage). Build, and fund it as a product team.
  • Regulatory or sovereignty constraints. Data residency, air-gapped deployments, FedRAMP High, or sector rules that no vendor meets. Buy is off the table; the question becomes self-host vs build.
  • No viable vendor exists at your requirements. If every vendor fails your 10× load test or lacks a critical semantic (e.g., exactly-once to your sink, per-tenant isolation), building is not vanity — it's the only option. Prove it with a benchmark, not an opinion.

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"We'll build our own, it's not that hard""I've never carried a pager for infrastructure""Building v1 is 20% of the cost. I'd price the 3-year on-call, upgrades, and staffing before committing."
"Open source is free, so self-host Kafka""I'm counting licenses, not people""Self-hosted Kafka is ~1–2 FTE of steady-state work. Below ~$1M/yr managed spend, MSK or Confluent is cheaper."
"We'll use DynamoDB, it scales infinitely""I haven't thought about exit cost or pricing model""DynamoDB is the right call for this access pattern; I'll keep it behind a repository interface and note that leaving costs ~N engineer-months."
"We can always migrate later""I don't understand data gravity""Swapping the SDK is cheap; moving 80 TB with dual writes is a two-quarter project. That makes this a one-way door, so I want more rigor now."
"The vendor has a 99.99% SLA""I think an SLA is a reliability mechanism""An SLA is a refund policy, not an architecture. I still need a degraded mode for when they're down."
"Let's abstract every vendor behind our own interface""I'll build a lowest-common-denominator layer nobody asked for""I'll put a seam where the exit cost is high and the vendor risk is real — not around every SDK."

Quick Reference#

Diagram: Quick Reference

Staff Sentence Templates#

"Before I pick a technology here, I want to ask whether [capability] is something our customers would notice if we did it better than competitors. If not, I'd buy it and spend our engineers on [actual differentiator]."

"Self-hosting [system] means a 24/7 rotation of 6–8 people and roughly [N] FTE of upgrade and tuning work. At our current spend of [$X/yr] on the managed version, the markup is cheaper than the team — I'd revisit when spend crosses [$threshold]."

"The exit cost here is roughly [N] engineer-months, driven by [data volume / call-site count / proprietary semantics]. That makes it a one-way door, so I'd put a [repository / adapter] seam at [boundary] now, while it costs a week instead of a quarter."

"I'd treat the vendor's SLA as a refund policy, not a reliability guarantee. When [vendor] is down, [feature] degrades to [fallback], and [team] owns that runbook."


Implementation Deep Dive#

1. The 3-Year TCO Model — Price People, Not Licenses#

The single most useful artifact in a build-vs-buy discussion is a TCO sheet with the people costs made explicit. Most bad decisions come from comparing a vendor invoice against "free" open source.

# All figures annualized; horizon = 3 years; FTE_COST = 300_000 (fully loaded)

function tco_managed(workload):
    usage     = workload.monthly_units * vendor.unit_price * 12
    growth    = compound(usage, workload.yoy_growth, years=3)
    discount  = negotiated_discount(commit=growth)          # 10-40% at committed spend
    glue_fte  = 0.25                                        # integration, IAM, cost monitoring
    exit_cost = estimate_exit(workload) * P(exit_in_3y)     # expected, not worst case
    return growth * (1 - discount) + glue_fte * FTE_COST * 3 + exit_cost

function tco_self_hosted(workload):
    infra     = compound(workload.raw_compute + storage + cross_az_traffic, growth, 3)
    ops_fte   = steady_state_fte(workload)                  # 0.5-2 per cluster family
    oncall    = max(0, ROTATION_MIN - existing_rotation) * FTE_COST * 0.15
    upgrades  = major_upgrades_in_3y * 0.5 * FTE_COST       # ~1 engineer-quarter each
    incidents = expected_sev2_per_year * hours_per_incident * blended_hourly * 3
    build     = initial_build_fte * FTE_COST                # usually 2-3x the estimate
    return infra + ops_fte * FTE_COST * 3 + oncall * 3 + upgrades + incidents + build

ROTATION_MIN = 6   # below 6 engineers, a 24/7 rotation burns people out

Worked Example — Kafka at Three Scales

ScaleManaged (MSK/Confluent-class) / yrSelf-hosted infra / yrSelf-hosted people / yrVerdict
10 MB/s, 3 brokers~$60–120K~$25–40K~0.75 FTE ≈ $225KBuy. People cost dwarfs the markup
200 MB/s, ~30 brokers~$0.8–1.5M~$300–500K~2 FTE ≈ $600KCoin flip. Decide on team availability and strategic value
2 GB/s, 300+ brokers~$6–12M~$2–3.5M~5 FTE ≈ $1.5MBuild/self-host if you can hire the team; the markup funds it several times over

Assumptions: $300K fully loaded FTE; public list-price order of magnitude; cross-AZ replication traffic included for self-hosted; negotiated discounts not applied. Real numbers vary 2× either way — the shape of the curve is the point.

🎯 Staff Insight: The crossover is not where infra cost equals managed cost — it's where infra + people equals managed cost. Because people cost is roughly a step function (you need a minimum viable team no matter how small the cluster), self-hosting is almost always wrong at small scale and often right at very large scale. The middle is where judgment lives.

2. The Differentiation Test — Is This Your Moat?#

Before any cost math, classify the capability. Cost is a tiebreaker; differentiation is the primary axis.

function classify(capability):
    score = 0
    if customers_would_notice_if_2x_better(capability):   score += 3
    if competitors_win_deals_on_it(capability):            score += 3
    if requirements_diverge_from_market(capability):       score += 2   # semantics no vendor offers
    if latency_or_cost_is_margin_critical(capability):     score += 2   # it's in COGS
    if vendor_market_is_mature(capability):                score -= 2   # 3+ credible vendors
    if talent_to_run_it_is_scarce_in_house(capability):    score -= 2

    if score >= 6:  return "CORE: build, staff as product"
    if score >= 3:  return "STRATEGIC: buy now, own the seam, revisit yearly"
    return "COMMODITY: buy, minimize customization"
CapabilityTypical ClassWhy
Payment ledger & reconciliationCore (fintech) / Strategic (others)Correctness is the product for a fintech; rails (card networks) are always bought
Identity / SSOCommodityCustomers expect it to work, never choose you because of it
Search rankingCore (marketplace, e-commerce)Relevance directly drives conversion
Search infrastructureCommodity / StrategicThe engine is commodity; the ranking model on top is the moat
Video delivery / CDNCore at Netflix scale, commodity otherwiseNetflix built Open Connect because delivery quality and cost are the product
ObservabilityCommodity → Strategic at scaleBills scale with cardinality; above ~$2–5M/yr many companies rebuild parts
Feature flagsCommodityMature vendor market; home-grown versions rarely add value
Diagram: 2. The Differentiation Test — Is This Your Moat?

🎯 Staff Insight: Split the capability before you classify it. "Search" is not one decision — the inverted-index engine is commodity, the relevance model is core. "Payments" is not one decision — card rails are bought, the ledger is built. The best build-vs-buy answers draw the line inside the capability.

3. Owning the Seam — Adapter Interface with a Migration Path#

Whether you buy or build, the highest-leverage move is a thin interface at the boundary that concentrates vendor-specific code in one place. This converts a 200-call-site migration into a one-adapter migration.

# Domain-level interface — named after YOUR semantics, not the vendor's
interface DocumentStore:
    put(tenant_id, doc_id, doc, if_version=None) -> Version
    get(tenant_id, doc_id) -> (doc, Version) | NotFound
    query_by_owner(tenant_id, owner_id, page_token) -> Page

class DynamoDocumentStore implements DocumentStore:
    # All DynamoDB-specific details live here: key design, GSIs, conditional writes
    def put(tenant_id, doc_id, doc, if_version):
        return ddb.put_item(
            Key = { pk: tenant_id + "#" + doc_id },
            Item = doc,
            ConditionExpression = "version = :v" if if_version else "attribute_not_exists(pk)")

# Migration: dual-write + shadow-read behind the same interface
class MigratingDocumentStore implements DocumentStore:
    def put(...):
        v = primary.put(...)                         # source of truth stays old system
        try: secondary.put(...)                      # best-effort, async-reconciled
        except e: metrics.increment("migration.dual_write.fail")
        return v
    def get(...):
        a = primary.get(...)
        if sample(1%): compare_async(a, secondary.get(...))   # "migration.shadow.mismatch"
        return a
Diagram: 3. Owning the Seam — Adapter Interface with a Migration Path

What the seam buys you: a 2-week adapter swap instead of a 6-month call-site hunt; a place to add metrics, retries, and circuit breakers consistently; and negotiating leverage — a credible exit is worth 10–30% at renewal.

What it costs: 10–20% more build effort up front, and the risk of a lowest-common-denominator API. The rule: name the interface after your domain semantics (conditional put, owner query), not a generic "KeyValueStore" that erases the features you chose the vendor for.

🎯 Staff Insight: Don't abstract the SDK — abstract the data contract. The expensive part of leaving DynamoDB isn't replacing put_item; it's that your code assumed single-digit-ms conditional writes and a 400 KB item limit. Put those assumptions in the interface's documentation so the exit estimate is honest.

4. The Open-Source Operational Ledger — What "Free" Actually Costs#

Self-hosting a stateful OSS system means inheriting a recurring workload that never appears in the adoption proposal. Track it explicitly.

Recurring WorkFrequencyEffort per OccurrenceAnnualized
Minor version / security patchesMonthly2–5 engineer-days (rolling restart, validation)~0.2 FTE
Major version upgrade (Kafka ZooKeeper→KRaft, ES 7→8, Cassandra 3→4)Every 12–24 months1–3 engineer-months~0.15–0.3 FTE
Capacity planning & rebalancingQuarterly1–2 engineer-weeks~0.1–0.15 FTE
Incident response (Sev2+)4–12 per year1–3 engineer-days incl. postmortem~0.1 FTE
Tuning (GC, compaction, JVM, OS)Continuous—~0.1–0.25 FTE
Backup / restore drillsQuarterly2–3 engineer-days~0.05 FTE
CVE response (e.g., Log4Shell-class)1–2 per year, unscheduledDays of all-handsUnbudgeted
Total per cluster family~0.7–1.5 FTE + on-call rotation share
# License-drift check — run yearly for every self-hosted OSS dependency
for dep in self_hosted_dependencies:
    assert dep.license in APPROVED_LICENSES            # Apache-2.0, MIT, BSD, PostgreSQL...
    assert dep.governance in ["foundation", "multi-vendor"] or dep.fork_plan_exists
    assert dep.version_lag_months <= 18                # older than this = security debt
    assert dep.named_owner and dep.oncall_rotation_size >= 6

License drift is a real risk class. Elastic moved Elasticsearch off Apache 2.0 in 2021 (AWS responded with the OpenSearch fork); HashiCorp moved Terraform to the BSL in 2023 (the community forked OpenTofu); Redis changed its license in 2024 (the Linux Foundation–backed Valkey fork followed). Single-vendor-governed OSS carries vendor risk even when you self-host it.

🎯 Staff Insight: Prefer OSS governed by a foundation (Apache, CNCF, Linux Foundation) for anything load-bearing. Single-company OSS is a vendor relationship with extra steps — the license can change under you, and your "free" option becomes a fork you have to maintain or a migration you didn't plan.

5. Vendor Due Diligence — Designing for the Vendor Being Down#

Buying doesn't outsource reliability; it outsources mechanism while you keep accountability. Before signing, run the vendor through the same questions you'd ask of an internal dependency.

QuestionGood AnswerRed Flag
What's the billing unit and its 10× driver?Predictable unit tied to revenue (requests, GB)Per-series, per-seat, or per-event pricing tied to engineering choices
Can we export all data in a documented format?Bulk export API, open formats, no egress penalty clause"Contact support"; proprietary formats; password hashes not exportable
Which regions, and what's the failure domain?Multi-region with published RTO/RPOSingle region, shared control plane across regions
What's the incident history?Public status page with 12+ months of postmortemsNo postmortems; vague "degraded performance" notices
What's the rate limit and quota model?Documented, per-tenant, raisable with noticeUndocumented; silent throttling
What happens on acquisition or sunset?Contractual notice period ≥ 12 months, data return clauseNone
Security and complianceSOC 2 Type II, ISO 27001, DPA, sub-processor listSelf-attested only
Diagram: 5. Vendor Due Diligence — Designing for the Vendor Being Down

The pattern: a timeout well below the caller's budget, a circuit breaker per vendor, an idempotency key so failover can't double-send, and a durable queue as the last resort. Wiring a secondary provider costs ~2 engineer-weeks and roughly 5–10% extra spend for a minimum commit; it converts a vendor outage from a Sev1 into a dashboard annotation.

🎯 Staff Insight: Decide per vendor whether you need redundancy (second provider) or graceful degradation (feature turns off). Email and SMS are cheap to make redundant. An identity provider or a payment processor is expensive to duplicate — there, invest in degradation: cached token validation, queued captures, and a clear customer-facing message.


Architecture Diagram#

Diagram: Architecture Diagram

Reading the diagram: product services never call vendors directly. Each seam is owned by the team that owns the capability class. The commodity piece (card rails, the search engine, message delivery) is bought; the differentiating piece (the ledger, the ranking model) is built. Notifications demonstrate vendor redundancy behind a seam — a second provider costs ~2 engineer-weeks to wire and turns a vendor outage from a Sev1 into a metric blip.


Failure Scenarios#

1. The Orphaned In-House Queue — Bus Factor of One#

Context: Three years ago, a strong engineer built a custom job queue on Postgres + Redis because "SQS didn't support priorities the way we wanted." It works. It processes ~40M jobs/day. The author left 8 months ago. Nobody on the current team has read the lease-renewal code.

t=0        Postgres primary fails over (routine, 25s)
t=+30s     Queue workers reconnect; lease-renewal code has a bug: it treats
           a connection reset as "lease still held" for 10 minutes
t=+2min    ~180K jobs held by dead leases; queue depth climbs 4K/sec
t=+15min   On-call finds no runbook; git blame points to a departed engineer
t=+45min   Someone manually expires leases with SQL; ~12K jobs execute twice
           (duplicate emails, 3 duplicate payouts caught by ledger reconciliation)
t=+3h      Backlog drained; postmortem opens

Detection: queue.depth rising while worker.active_jobs flat; queue.lease.age_p99 > 5 min (normal: 30s); jobs.duplicate_execution from downstream idempotency checks.

Blast radius: Every async workflow in the company — 23 services enqueue into it. Duplicate payouts required finance involvement.

Mitigation: Manual lease expiry, then pause non-critical producers until drained.

Prevention: (1) Assign a named owning team and a rotation of ≥6; (2) write the failover runbook and run it as a game day quarterly; (3) re-run the build-vs-buy analysis — SQS FIFO + a priority-topic scheme or a managed workflow engine now covers the original requirement.

Owner: Infrastructure platform lead, with the VP Eng signing off on "adopt and staff" vs "migrate off."

🎯 Staff Insight: Every home-grown system eventually becomes an orphan unless someone is funded to own it. The build decision must include the ownership decision: which team, what rotation, what happens when the author leaves. "We'll figure it out" is how a clever v1 becomes a Sev1 three years later.

2. The Observability Bill That Tripled at Renewal#

Context: A SaaS observability vendor was adopted at $25K/month. Two years later, a Kubernetes migration multiplied pod-level tags, and a product team added user_id as a metric tag.

Month 0     Bill: $25K/mo, 400K active custom metric series
Month 14    K8s migration: pod names in tags, series count 2.1M
Month 20    user_id tag added to 6 metrics: series count 9M
Month 22    Renewal quote: $310K/mo (list), $190K/mo (negotiated)
Month 22    CFO escalates; eng has 60 days to choose: pay, cut, or leave

Detection: Should have been vendor.billable_units tracked daily with a budget alert at +20% MoM; instead it was discovered on an invoice.

Blast radius: Organizational, not technical — a $2M/yr unbudgeted line item, and an emergency migration project that pulls 4 engineers off roadmap for a quarter.

Mitigation: Emergency cardinality controls (drop user_id tags, aggregate pod-level series to deployment level), cutting series ~70% within 3 weeks.

Prevention: Cost as an SLO: per-team showback of vendor spend, a cardinality budget enforced at the collector, and a seam (OpenTelemetry collectors) so the vendor becomes swappable. The seam alone was worth ~25% at the next negotiation.

Owner: Observability platform team owns the collector and cardinality budget; each product team owns its showback line.

🎯 Staff Insight: SaaS pricing scales with your behavior, often super-linearly — per host, per series, per event, per seat. Before buying, ask: "What's the billing unit, and which engineering decision multiplies it by 10×?" For observability it's tag cardinality; for auth it's MAUs; for search SaaS it's records × operations.

3. Correlated Vendor Outage — The Identity Provider Takes Down 14 Services#

Context: All customer-facing and internal login flows use one managed identity provider in one region. Fourteen teams adopted it independently; nobody tracked the aggregate dependency.

t=0        IdP has a regional incident: token issuance p99 goes 80ms -> 30s timeouts
t=+1min    Web and mobile login fail; existing sessions still valid (1h tokens)
t=+5min    Internal admin tools fail SSO; support can't access customer records
t=+40min   Short-lived service-to-service tokens (15 min TTL) begin expiring
t=+55min   Two backend services fail authz checks; checkout error rate 18%
t=+2h      IdP recovers; token refresh storm (8x normal) causes a second 10-min brownout

Detection: auth.token_issue.latency_p99 and auth.token_issue.error_rate on the seam; synthetic login probes every 30s from 3 regions.

Blast radius: All new logins company-wide; checkout after token expiry; internal support tooling. Revenue impact from the checkout failures.

Mitigation: Extend service-token TTLs via config; enable cached-JWKS validation; serve "logged-in users continue, new logins queued" banner.

Prevention: (1) Separate service-to-service auth from end-user IdP (workload identity owned internally); (2) cache signing keys and validate tokens locally so validation never calls the vendor; (3) jittered refresh; (4) an org-level registry of critical vendors with a concentration review.

Owner: Identity platform team; the concentration registry is owned by the architecture review group.

🧭 Principal Insight: No single team made a bad decision — each of 14 teams rationally bought the same well-reviewed vendor. The failure was organizational: nobody owned the aggregate dependency. Correlated vendor risk is invisible at the team level and only fixable at the org level.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Home-grown system orphanedownership.unassigned_services, runbook age > 12 monthsEvery consumer of the systemAssign team, fund rotation, or plan migrationVP Eng / platform lead
SaaS bill spikevendor.billable_units MoM +20%Budget; roadmap if emergency migration neededCardinality/usage caps, negotiate, activate seamPlatform team + finance partner
Vendor outageSeam-level vendor.error_rate, synthetic probesAll dependents of that vendorDegraded mode, secondary provider, cached credentialsCapability-owning team
License change on self-hosted OSSDependency license scan, foundation watchEvery team self-hosting itPin last permissive version, move to fork, or buy managedArchitecture review / OSPO
Managed service version lagdep.version_lag_months > 18Security posture; blocked featuresSchedule provider upgrade, escalate via TAMPlatform team
Vendor acquired / sunsetVendor viability review, contract noticesEvery integrationExecute exit plan behind seamProcurement + capability owner

The Principal Lens#

Why L7 Sees This Problem Differently#

A Staff engineer makes a good build-vs-buy decision for one system. A Principal engineer recognizes that the org makes dozens of these decisions a year, mostly implicitly, and that their sum defines the company's engineering cost structure, risk posture, and where its best people spend their time. The L7 job is not to pick Kafka vs SQS — it's to establish the decision rights, the defaults, and the review threshold so 40 teams make consistent choices without a Principal in the room, and to periodically retire capabilities the company should never have built or should no longer buy.

The Org-Level Fault Line#

Team autonomy vs portfolio coherence. Letting every team pick its own vendors and self-host what it likes maximizes local speed and produces 4 message brokers, 3 feature-flag systems, 2 observability vendors, and zero negotiating leverage. Centralizing every choice produces a platform team bottleneck and a 6-week architecture review for a logging library. The Principal position: centralize the capability classes where consolidation creates leverage (data stores, messaging, identity, observability, CI) with a paved road and a lightweight exception process; leave everything else to teams with only a spend threshold and a security review.

Cost Model#

ScaleExample Capability PortfolioManaged-Heavy ApproachBuild/Self-Host-Heavy ApproachAssumptions
Startup (30 eng, $5M ARR)DB, queue, search, auth, observability~$25–60K/mo vendors; ~1 FTE platform~$10–20K/mo infra; ~4–5 FTE to run it allSelf-hosting here spends 15% of eng on non-differentiating work
Growth (300 eng, $150M ARR)+ streaming, workflow, feature flags~$0.8–1.5M/mo; ~8 FTE platform~$0.4–0.7M/mo infra; ~25 FTE platformMix: buy most, self-host 1–2 where spend > $1M/yr
Large (3,000 eng, $3B ARR)Full portfolio, multi-region~$10–20M/mo~$5–9M/mo infra; ~150 FTE platform orgRepatriating 2–3 largest line items typically saves $20–60M/yr net

Assumptions: $300K fully loaded FTE; vendor spend at list order of magnitude with 20–35% enterprise discount at the large tier; on-call rotations of 6–8 per owned capability.

🧭 Principal Move: "I don't want 40 separate build-vs-buy decisions. I want a paved-road list: here are the 8 capability classes we centralize, here's the default for each, and any team deviating needs to show a 3-year TCO and a named on-call owner. Everything off that list, teams decide themselves."

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to ReverseWhy
Email / SMS delivery vendorTwo-way~2–4 engineer-weeksStateless, narrow API, easy to run two in parallel
Feature-flag vendorTwo-way~1–2 engineer-monthsMany call sites, but flag state is small and exportable
Observability vendor (with OTel seam)Two-way-ish~1 engineer-quarterDashboards and alerts must be rebuilt; data is disposable after retention
Identity provider (end-user)Mostly one-way~2–4 engineer-quartersPassword hashes may not be exportable; forced resets hurt users
Proprietary primary datastore (DynamoDB, Spanner, Firestore)One-way~2–6 engineer-quarters per major serviceData gravity + semantics baked into code (conditional writes, item limits, consistency)
Cloud provider (primary)One-wayYears, $10M+ at scaleEvery service, IAM model, network topology, and team skill set
Building a core capability in-houseOne-way-ishHeadcount + morale; killing a team's system is hardSunk-cost bias makes teams defend what they built

🧭 Principal Insight: The most underrated one-way door is building, not buying. Vendors can be fired; an internal team that has spent three years on a system will fight its retirement. Budget the retirement conversation into the build decision.

The Standard I'd Write#

RFC: Build, Buy, and Self-Host Decisions for Shared Capabilities

Scope: Any new infrastructure capability or vendor with projected spend > $100K/yr, or any self-hosted stateful system requiring on-call.

Mandatory requirements:

  • Decisions MUST include a 3-year TCO including fully loaded people cost, on-call rotation, and expected exit cost.
  • Self-hosted or built systems MUST have a named owning team with an on-call rotation of ≥ 6 engineers before production traffic.
  • Vendors in a paved-road capability class (data stores, messaging, identity, observability, CI/CD, workflow, secrets, feature flags) MUST use the paved-road default unless an exception is approved.
  • Integrations with projected exit cost > 3 engineer-months MUST go through an owned internal interface.
  • Self-hosted OSS SHOULD be foundation-governed; single-vendor-governed OSS requires a documented fork/exit plan.
  • Every vendor MUST have per-team spend showback and a budget alert at +20% MoM.

Exceptions: Filed as a one-page doc to the architecture review group; decision within 10 business days; exceptions expire after 12 months and must be renewed.

Success metrics: Duplicate vendors per capability class ≤ 1; orphaned systems (no owner) = 0; vendor spend growth ≤ revenue growth; median exception turnaround ≤ 10 days.

What I'd Tell the VP#

We are spending roughly a third of our platform engineers running systems our customers will never notice, and we have no view of which vendors, if they went down, would take out the whole company. I want to do three things. First, publish a short list of default choices for eight infrastructure categories so teams stop re-deciding them. Second, move our two largest, most predictable cloud bills to self-run infrastructure where the savings pay for the team several times over. Third, review once a year what we built that we should stop running. Expected result: lower infrastructure growth than revenue growth, fewer company-wide outages from a single vendor, and more engineers on product work.

Principal Interview Signals#

SignalWhat It Sounds Like
Portfolio thinking"This isn't one decision — 12 teams will hit this. I'd make it a paved-road default with an exception path."
Prices opportunity cost"The real cost of self-hosting isn't the 2 FTE — it's the 2 FTE not shipping the matching algorithm we compete on."
Correlated risk"Before we add another dependency on this vendor, how many tier-0 services already share its fate?"
Deliberate lock-in"We accept lock-in to one cloud for the 30% discount and the single ops model. We hedge observability and identity, where exit cost is still low."
Retirement discipline"Every year we should ask what we built that a vendor now does better, and kill it — that's harder than building."

Staff answers that L7 interviewers find insufficient:

  • "I'd compute the TCO and pick the cheaper option" — correct for one system, silent on the 40 teams making the same choice independently.
  • "I'd abstract the vendor behind an interface" — a good local move, but no view on which capabilities warrant the abstraction tax org-wide.
  • "The platform team will own it" — without checking whether that team has the headcount, or whether its other commitments get worse.

In the Wild#

These are publicly documented decisions — each is famous precisely because it went against the default in a specific, defensible way.

Dropbox: Magic Pocket — Building the Storage Layer That Was the Product#

Dropbox launched on Amazon S3 and ran on it for years. Around 2015–2016 it migrated the large majority of user file data — hundreds of petabytes — onto Magic Pocket, an in-house exabyte-scale block storage system on custom hardware. Its 2018 S-1 filing attributed roughly $75M of infrastructure savings over two years to the move. Crucially, Dropbox kept using AWS for parts of the business and for some regions; this was a targeted repatriation of the single largest, most predictable line item, not a blanket "leave the cloud."

Staff insight: The decision worked because every condition for "build" lined up: storage was the product and the dominant cost of goods sold, the workload was enormous and predictable, and Dropbox could fund a dedicated team of storage specialists. In an interview, cite this as the exception that proves the rule — and name the conditions, not just the savings.

Netflix: Buy the Cloud, Build the CDN#

Netflix spent roughly seven years (2008–2016) moving its streaming infrastructure from its own data centers onto AWS — a decisive "buy" for compute, storage, and databases. At the same time, it built Open Connect, its own content delivery network, placing Netflix-designed caching appliances inside ISP networks. Video delivery quality and bandwidth cost are the heart of the streaming business; generic compute is not.

Staff insight: Netflix split the stack by differentiation, not by ideology. The same company made opposite decisions for two layers in the same years. That's the answer shape interviewers want: "buy the commodity layer, build the layer where our customers feel the difference."

37signals: Cloud Repatriation at Modest Scale#

37signals (Basecamp, HEY) publicly documented leaving AWS for its own hardware in colocation starting in 2022–2023, reporting that its roughly $3M+/yr cloud bill could be replaced with owned servers, with projected savings in the millions over five years. The workload was mature and stable, and the company already had experienced operations engineers.

Staff insight: Repatriation isn't only for hyperscalers — but the enabling conditions are specific: flat, predictable load; an existing ops team (so the people cost is already sunk); and low reliance on proprietary managed services. A company with 40 teams on DynamoDB, Step Functions, and Lambda could not replicate this cheaply — its exit cost is in the code, not the servers.


Practice Drill#

Prompt: "Your company runs ~150 microservices. The Datadog bill is $220K/month and growing 8% month over month. The CFO asks whether you should build your own observability stack on Prometheus/Grafana-class open source. What do you recommend?"

Staff Answer

I'd start by separating why the bill grows from whether to leave. 8% MoM is ~2.5× per year — that's almost always a cardinality or ingest problem (per-pod tags, high-cardinality labels, debug logs at INFO), not traffic. Step one is two to three weeks of cost hygiene: a cardinality budget at the collector, drop or aggregate the top 20 series families, sample traces at 10% for healthy paths, and route logs by tier. That typically cuts 30–50%. Step two is putting an OpenTelemetry collector layer in front of the vendor — a seam that makes us portable and gives us negotiating leverage. Only then do I compare TCO: at a post-hygiene $130K/month ($1.5M/yr), a self-run metrics + logs + traces stack costs $400–600K/yr infra plus a 5–6 person observability team ($1.7M/yr) — roughly break-even, with significant migration risk and a year of degraded tooling. So my recommendation is: cut cost now, build the seam, renegotiate with a credible exit, and move only the highest-volume, lowest-value signal (e.g., raw logs retained for compliance) to cheap self-run storage. Revisit a full move if spend crosses ~$3M/yr after hygiene.

Why this is L6:

  • Diagnoses the growth driver before deciding — the bill is a symptom
  • Prices the self-hosted team (5–6 FTE) instead of treating OSS as free
  • Uses a seam to turn a one-way door into a two-way door and to gain leverage
  • Splits the capability (logs vs metrics vs traces) instead of deciding it wholesale

What L7 adds:

  • Sets org-wide cardinality and log-tier policy with per-team showback, so the bill can't silently regrow after the fix
  • Frames the vendor contract as a multi-year portfolio negotiation (commit level, price protection, exit terms) alongside procurement
  • Names the observability platform as a funded team with a charter, rather than a side job for SREs

Prompt: "A team wants to build its own workflow engine because Temporal 'is too heavy.' How do you respond?"

Staff Answer

I'd ask what "too heavy" means concretely — operating a Temporal cluster (Cassandra/Postgres + history/matching services) is real work, roughly 1–1.5 FTE self-hosted. But a home-grown workflow engine must solve durable state, timers, retries with backoff, idempotent activity dispatch, versioning of in-flight workflows, and visibility — 12–18 engineer-months to reach parity with the parts they'll need, then a permanent owner. The options I'd compare: Temporal Cloud or AWS Step Functions (buy, near-zero ops), self-hosted Temporal (moderate ops), or a simpler async job queue if the workflows are really just 2–3 step jobs. If the workflows are short and linear, a queue + state table is the right answer and "build" is fine because it's small. If they're long-running with compensation logic, building a workflow engine is building a database — I'd buy.

Why this is L6:

  • Turns "too heavy" into specific operational costs
  • Enumerates the hidden requirements a home-grown engine must meet
  • Matches solution weight to actual workflow complexity

What L7 adds:

  • Checks whether other teams already run a workflow engine and makes one the paved road
  • Sets a policy: new workflow engines require architecture review; existing one-offs get a migration timeline

Staff Interview Application#

How to Introduce This Pattern#

"Before I choose between building this and using [vendor/OSS], I want to classify it. If it's not something our customers would notice us doing better, I'd default to a managed service and spend engineers on [differentiator]. I'll sanity-check that with a rough 3-year TCO — including the on-call rotation, which is the part that usually flips the answer — and I'll put a thin interface at the boundary so the decision stays reversible."

Lead with differentiation, then cost, then reversibility. That order signals that you think about the business first and the technology second.

When NOT to Use This Pattern#

  • The decision is trivially cheap to reverse. Choosing a JSON library or an email vendor doesn't need a TCO spreadsheet. Pick, move on, and spend rigor on one-way doors.
  • The interviewer has already fixed the stack. If the prompt says "you're on AWS with DynamoDB," don't relitigate — design within it and mention exit cost only if it changes a decision.
  • Prototype or throwaway scope. For a 3-month experiment, buy everything and optimize nothing; the TCO horizon is shorter than the contract.
  • Regulation already decided it. If data can't leave your facility, "buy SaaS" is off the table; the analysis is only self-host vs build.

Follow-Up Questions to Anticipate#

Interviewer AsksWhat They Are TestingHow to Respond
"Wouldn't it be cheaper to self-host?"Whether you count people"Infra would be ~40% cheaper; with a 6-person rotation and upgrade work, it's more expensive until spend passes roughly $1–2M/yr."
"What if the vendor goes down?"Degraded-mode design"The SLA is a refund policy. Behind the seam I'd have [fallback: cached tokens / secondary provider / queue-and-retry], and the owning team has the runbook."
"How would you get off this vendor?"Exit cost awareness"The seam limits code change to one adapter; the real cost is data — [N TB] with dual-write and shadow-read, roughly [N] engineer-months."
"Why not build it — we have great engineers?"Opportunity cost"We do, which is exactly why I'd rather they build [differentiator]. Building this well means staffing it for years, not months."
"Open source vs managed version of the same thing?"Operational burden"Same API, so the exit is cheap either way. I'd start managed and self-host only if spend and team capacity justify it."
"What if the OSS license changes?"Governance risk"I prefer foundation-governed projects. For single-vendor OSS, I pin a version and know which fork I'd move to — Elastic, HashiCorp, and Redis all changed licenses in the last few years."

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
FramingFeatures and benchmarksDifferentiation first, then TCOPortfolio: which capabilities the org should own at all
CostLicense vs "free"3-yr TCO with FTEs and on-callOpportunity cost, vendor concentration, negotiation leverage
RiskTrusts SLADesigns degraded mode; prices exitCorrelated vendor risk across teams; license and viability risk
Ownership"Someone will run it"Named team, rotation ≥ 6, runbookDecision rights, paved road, exception process, retirement review

Strong hire signals: splits a capability into commodity and core parts; volunteers the on-call cost unprompted; quantifies exit cost in engineer-months; names the conditions under which they'd reverse the decision.

Lean no-hire signals: "open source is free"; "we'll build it, it's a weekend project"; "we can always migrate later" with no estimate; treating an SLA as an availability guarantee.

Common false positive: Deep knowledge of Kafka broker tuning ≠ good build-vs-buy judgment. Candidates who know a system intimately often overvalue self-hosting it because the operational cost feels small to them personally — it isn't small for the team that inherits it.


Capacity Planning Quick Reference#

NumberValueContext
Fully loaded engineer~$250–400K/yrSalary + benefits + equity + overhead; use $300K as the default
Sustainable 24/7 rotation6–8 engineersBelow 6, on-call frequency exceeds 1 week in 5 and attrition follows
Steady-state ops per stateful OSS cluster family0.5–2 FTEKafka, Cassandra, Elasticsearch, self-hosted Postgres fleets
Major version upgrade1–3 engineer-monthsEvery 12–24 months per system
Managed markup over raw compute~2–4×Varies by service; storage-heavy services often lower
Repatriation crossover~$1–2M/yr managed spend, stable loadPlus an available team; spiky loads shift it higher
Enterprise commit discount~10–40%Proportional to commitment term and volume
Seam / adapter overhead10–20% extra build effortWorth it when exit cost > ~3 engineer-months
Home-grown build overrun2–3× initial estimateBefore the ongoing ownership cost
Vendor spend alert+20% MoMCatch cardinality/usage explosions before renewal

Pitfalls checklist:

  • Did we count on-call and upgrade people-cost, not just infra?
  • Is there a named owner with a rotation of ≥ 6 before launch?
  • What's the billing unit, and which engineering decision multiplies it 10×?
  • What's the exit cost in engineer-months, and does it justify a seam?
  • Is the OSS foundation-governed, and what's the fork plan if not?
  • How many other tier-0 services already depend on this vendor?
  • When will we revisit this decision, and what metric triggers it?
  1. Loading the index…