Hiring BarSupport

Design a Calendar and Scheduling Service

Case study103 min read10 diagrams

Technologies referenced in this case study: PostgreSQL · Cassandra · Redis · Kafka · Distributed SQL

Related: Job Scheduler · Notifications · Reservation Systems · Cloud File Sync · Schema Design · Hot Keys · Real-Time Updates · Idempotency

Reading Guide#

Organized for interview use first, reference second. This page designs a calendar service: events, recurring series, invitations and RSVPs, free/busy across many people, rooms, booking pages, reminders, and sync with clients the service does not control. Three neighbours own pieces of the machinery and this page links to them instead of repeating them: firing a timer reliably at a wall-clock moment is Job Scheduler; getting a push or email to a phone is Notifications; never selling one room twice is Reservation Systems. Cursor-based sync with offline clients is covered in depth in Cloud File Sync.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on recurrence expansion, the time model and sync
What is a Calendar & Scheduling Service? — Why interviewers pick this topic

A calendar service stores commitments in civil time — "every Tuesday at 09:30 in London", "the board meeting on the last Thursday of each quarter", "Priya's birthday on 29 February" — and answers three questions about them, millions of times a second: what is on my calendar this week, when are these eight people and one room all free, and what should ring on my phone in ten minutes. On top of that it moves invitations between people (inside the company and outside it), keeps phones, laptops and third-party clients in sync, and lets strangers book a slot on a booking page.

The hard part is not storing a row with a start and an end. The hard part is that most of the events users care about have not happened yet, and the meaning of a future time can change after you store it. A weekly meeting has no last instance. "09:30 London" maps to a different UTC instant in July than in January. Governments change daylight-saving rules with days of notice, and every future meeting in that zone must move — or must not move, depending on which zone the organizer meant. One edit to a series with 300 attendees touches 300 calendars, thousands of reminder timers, and every device those people own.

Before vs After — the "government changes the clocks" scenario:

Without a designed time model:
t=-2y:      Events stored as UTC instants. A weekly 09:00 Mexico City stand-up is
            expanded into 104 rows of UTC timestamps, assuming DST continues.
t=0:        Mexico abolishes DST for most of the country. New tz data ships.
t=+1day:    Nothing recomputes. Stored instants are now one hour off for every
            winter instance. Phones (with new OS tz data) render 08:00 for some
            users and 09:00 for others, depending on whether the client re-derives.
t=+1week:   First Monday after the change: half the team joins an hour early.
            Room bookings and free/busy now disagree with what users see.
t=+3weeks:  Support finds 4M affected series. Fix requires a migration over
            data whose original intent ("09:00 Mexico City") was never stored.

With a designed time model:
t=0:        Events stored as (local start 09:00, zone America/Mexico_City, RRULE).
            Derived UTC instants in the index carry the tzdb version used.
t=+1h:      New tzdb release ingested; diff shows which zones changed from which date.
t=+2h:      Rebase job recomputes future index entries and reminders for affected
            zones only. Past instances are untouched.
t=+3h:      Every client and the free/busy index agree: 09:00 local, new offset.
            Cross-zone attendees see the meeting move in their own zone — correctly.

Why interviewers reach for this question: It looks like CRUD with a nice UI, so it cleanly separates candidates who model rows from candidates who model intent. The traps are all correctness traps that pass every demo: storing UTC for future events, materializing infinite series, treating the organizer's and the attendee's copies as the same object, computing free/busy by scanning events, and scheduling reminders once at creation time. Each one fails months later, on a DST weekend or a policy change, for millions of users at once.

Mechanics Refresher: The Calendar Primitives
PrimitiveHow It WorksProsCons
UTC instant2026-11-03T14:00:00ZUnambiguous; sorts and compares triviallyLoses intent: "09:00 in Chicago" cannot be recovered after a rule change
Wall-clock + IANA zone09:00 + America/ChicagoPreserves what the user meant; survives tz rule changesMust be converted to compare; DST gaps and overlaps need a policy
Floating time09:00 with no zone"Take medication at 08:00 wherever I am"Means a different instant for every viewer; cannot go in a shared free/busy index
All-day / date value2026-12-25 (a date, not an instant)Holidays and birthdays don't drift across zonesSpans different UTC ranges per viewer; must not be stored as midnight UTC
RRULEFREQ=WEEKLY;BYDAY=TU;INTERVAL=1 from RFC 5545One row describes an infinite seriesMust be expanded to answer any time-range question
Exception (RECURRENCE-ID)An override keyed by the instance's original startMoves or edits one instance without touching the ruleOrphaned when the rule changes underneath it
EXDATEList of cancelled original startsCancels single instances cheaplySame orphaning risk; lists grow over years
Series split ("this and following")End the old rule with UNTIL, start a new seriesClean history; past instances keep old detailsTwo series to keep linked; exceptions must be re-homed
Free/busyBusy intervals only, no titlesPrivacy-preserving availability across peopleMust reflect expanded recurrences and every RSVP
iTIP messageREQUEST / REPLY / CANCEL with UID + SEQUENCE (RFC 5546)Interoperable invites across systemsOrdering and duplicates are the receiver's problem
Sync tokenOpaque cursor into a calendar's change logClients fetch only deltasExpiry forces a full resync; a storm if many expire at once

For most production systems: store future events as wall-clock + IANA zone + RRULE + exceptions (the intent), derive UTC instants into a bounded-horizon index stamped with the tzdb version (the cache), keep one shared event with per-attendee overlay rows inside the system and iTIP at the boundary, answer free/busy from the busy-interval index, and materialize reminder timers on a rolling horizon so edits and tz changes re-derive them. The primitives are not the interview — what each stored value means when the rules change is.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A calendar is not a CRUD question. Anyone can store an event with a start and an end.

It is a time-and-recurrence correctness question that tests:

  • Whether you store what the user meant (a wall-clock time in a named zone, and a rule) rather than what it happened to mean on the day you stored it
  • Whether you can answer "what is on these calendars between T1 and T2" when most events are infinite rules plus exceptions
  • Whether you see that one event has many owners — the organizer owns the time, each attendee owns their RSVP, reminders and colour — and design the write path around that split
  • Whether you treat reminders, free/busy and device sync as derived views that must be re-derived when the rule, the RSVP or the tz database changes

The key insight: For a calendar, the source of truth for a future event is intent, and every instant is a cache. Past events are facts in UTC; future events are wall-clock times in a zone, governed by rules that both users and governments can change. Staff candidates say this in the first five minutes and then design the free/busy index, the reminder timers and the sync change log as re-derivable from intent — with an explicit job that re-derives them when the tz database changes.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws events table → API → clients; adds recurring events as a flagAsks "Personal, enterprise scheduling or booking pages? Which zone anchors a meeting? How far ahead must free/busy be exact, and which clients do we not control?"Asks "How many time libraries and tz data copies do our server, web, iOS, Android and email stacks embed, and how do we know they agree?"
Time"Store everything in UTC, convert on display""UTC for the past, wall-clock + IANA zone for the future. Derived instants carry a tzdb version; a tz release triggers a rebase of affected future instances."Owns tz data as a company-wide dependency: one ingestion pipeline, a skew SLO across clients, and a short-notice-change runbook
Recurrence"Generate all instances into the events table""Store the RRULE and exceptions keyed by original start; expand on read; materialize only a bounded horizon into the busy index"Publishes one recurrence engine with a conformance suite every client must pass; RRULE semantics become a one-way-door contract
Invites"Copy the event into each attendee's calendar""One shared event owned by the organizer, a per-attendee overlay for RSVP and reminders, iTIP with SEQUENCE at the boundary for external attendees"Decides interop posture: which external protocols (CalDAV, iMIP, Exchange) are first-class, and who funds their long tail
Free/busy"Query each attendee's events for the range""A per-calendar busy-interval index, 18 months materialized, merged per query; privacy enforced by returning intervals only"Treats availability as a product API with a published freshness SLO that rooms, booking pages and partner tools build on
Scale"Shard events by user""Shard by calendar; the hot reads are sync 'anything changed?' checks (~450K/s) and the hot spike is reminders at :50 past the hour"Prices it: sync and reminder fan-out dominate cost; decides what runs on devices (local alarms, expansion) to cut server load
Why "time" separates levels

L5: "Everything is stored in UTC; the client converts to local time for display." This is correct for logs, payments and past events, and it is the standard advice for most backends. For a future recurring meeting it silently destroys information: once "09:00 in Chicago every Tuesday" is turned into a list of UTC instants, a change to Chicago's DST rules cannot be applied, and you cannot even tell which events were meant to follow Chicago rather than the UTC offset they happened to have.

L6: "A future event is stored as a local start, a duration and an IANA zone ID — America/Chicago, never CST or -06:00. I derive UTC instants into the index for range queries, and every derived row records the tzdb version it was computed with. When a new tzdb release lands, I diff it, find zones whose future offsets changed, and recompute only those future instances, their busy intervals and their reminders. Past instances freeze as UTC: what happened, happened. All-day events are dates, and 'take my pill at 08:00' is floating time with no zone at all."

L7: "Our server uses one tz library, the web client another, iOS and Android use whatever the OS ships, and the email renderer has its own. After a short-notice change those disagree for weeks. I'd make tz data a platform dependency with an owner, an ingestion SLO measured in hours, and a dashboard of tzdb versions observed across clients — because the bug users see is not 'wrong time', it's 'my phone and my laptop disagree'."

Why "recurrence" separates levels

L5: "When someone creates a weekly meeting, we insert a row per instance for the next two years." Reasonable for a demo. Then the organizer changes the time: 104 rows rewritten, exceptions lost, and the series silently ends two years from creation. A daily standup with 40 attendees materialized for two years is ~29,000 attendee-instance rows from one click.

L6: "The series is one row with an RRULE. Exceptions are separate rows keyed by (series_id, original_start) — the RFC 5545 RECURRENCE-ID — so they survive edits to other instances. Range queries expand the rule on read; it's microseconds per instance. The only thing I materialize is busy intervals for the next 18 months, because free/busy across 20 people needs an index, not 20 rule expansions — and that materialization is a cache I can rebuild."

L7: "There are five RRULE expanders in this company — server, web, two mobile apps and the email digest — and each disagrees with the others on at least one edge case: BYSETPOS, the 31st of short months, DST gaps. I'd ship one engine, compiled for every platform or behind one service, with a shared conformance suite, and I'd treat its semantics as a contract we never change silently."

Why "invites" separate levels

L5: "Each attendee gets their own copy of the event in their calendar." It makes reads simple and it makes every edit a fan-out write to N copies that can disagree. A 300-person all-hands moved by 30 minutes is 300 writes, 300 sync notifications and, if one copy fails, one person who shows up at the old time.

L6: "Inside our system there is one event, owned by the organizer's calendar. Each attendee has an overlay row: RSVP status, personal reminders, colour, whether it's hidden. Time, title and rule live only on the event, so an edit is one write plus index and notification fan-out. External attendees get iTIP messages over email — REQUEST, REPLY, CANCEL — and their SEQUENCE numbers tell both sides which version wins."

L7: "Interop is where this product's cost hides: Exchange, CalDAV servers, ICS feeds, mail clients that render invites their own way. I'd decide explicitly which protocols get first-class support and an owning team, and which are best-effort — because otherwise every enterprise deal adds one more half-supported integration."

Positions to Commit To#

PositionRationale
Future events are wall-clock + IANA zone; past events are UTC instantsGovernments change rules with days of notice; only the intent survives that. The past cannot change, so freeze it
Store the rule, expand on read, materialize a bounded horizonInfinite series cannot be rows; free/busy needs an index; 18 months covers almost all scheduling while staying rebuildable
Exceptions are keyed by original start, never by current startRECURRENCE-ID is what keeps a moved instance attached to its series across edits
One shared event, per-attendee overlays; iTIP only at the boundaryAn edit is one write, not N; RSVPs never contend with the organizer's edits
Free/busy is a derived busy-interval index, not a query over events8 attendees × rule expansion per query does not survive 80K queries/s; the index does
Reminders are a rolling-horizon projection, re-derived on every changeScheduling at creation time goes stale on the first edit, RSVP decline or tz release
Sync is a per-calendar change log with tokens; push is a hint, poll is the guaranteeClients you don't control miss pushes; tokens that expire must degrade to a resync, not data loss

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Personal & shared calendars (consumer)500M users, many devices each, offline edits, third-party clientsIntent-based time model, expand-on-read, change-log sync, device-local alarms plus server remindersWrong time after a DST or tz rule change; devices disagree; missed or duplicate remindersEvery client renders the same local time for an instance; reminder on time p99 ≤ 30 s
Enterprise scheduling (free/busy, rooms, delegation)60M seats, meetings across zones and companies, rooms as contended resources, privacy rulesBusy-interval index, organizer-owned events with attendee overlays, room booking as a conditional write, iTIP with external systemsDouble-booked rooms; free/busy stale after edits; private details leaking through availabilityFree/busy reflects writes within 5 s; zero double-booked rooms; no title leaks to non-delegates
Booking pages / appointment schedulingStrangers book slots in a host's availability; bursts when a link is shared; payments sometimes attachedAvailability = working hours − busy − buffers, computed from the same index; booking as an idempotent conditional insertTwo bookers get the same slot; host's new meeting races a booking; slot shown in the wrong zoneZero double bookings; slot times shown in the booker's zone and stored in the host's

🎯 Staff Move: "I'll design the shared core first — the time model, recurrence, the organizer-owned event and the busy-interval index — because personal calendars and enterprise scheduling both stand on it. I'll go deepest on enterprise scheduling: free/busy across many people, rooms and invites are where the correctness bar is highest. Booking pages reuse the same availability computation with a reservation-style write path, which I'll cover at the end."

Where the Design Splits#

#Fault LineThe Tension
1Recurrence: Store the Rule vs Materialize InstancesCompact, editable, infinite series vs simple range queries and indexable instances
2Time Representation: UTC Instant vs Wall-Clock + ZoneTrivial comparison and sorting vs surviving tz rule changes and DST without losing intent
3Invitations: Shared Event vs Per-Attendee CopiesOne write per edit and one truth vs independent copies that read cheaply and work across systems
4Free/Busy: Compute on Read vs Precomputed Busy IndexAlways-correct, no extra state vs fast multi-person queries with an index that can go stale
5Reminders: Schedule at Write vs Rolling-Horizon ProjectionSimple one-time timers vs reminders that follow edits, RSVPs and tz changes

How Real Companies Built It#

Why this section belongs here: Calendars run on public standards and public data. Every fault line on this page is visible in an RFC, the tz database's release history, or a large provider's documented API behavior.

RFC 5545 — iCalendar Recurrence Has Sharp Edges by Design#

RFC 5545 (September 2009) defines the RRULE grammar every major calendar speaks, and it settles three edge cases on paper: an RRULE that generates an invalid date (30 February) or a nonexistent local time (inside a DST spring-forward gap) MUST ignore that instance and not count it toward COUNT; a DATE-TIME whose local time occurs twice in a fall-back overlap refers to the first occurrence; and a DTSTART inside the gap is interpreted with the offset from before the gap, so America/New_York 02:30 on 11 March 2007 means 03:30 EDT. It also defines RECURRENCE-ID with RANGE=THISANDFUTURE, and forbids COUNT and UNTIL in the same rule (RFC 5545). RFC 7529 (May 2015) later added RSCALE for non-Gregorian calendars and a SKIP option — OMIT (the default), BACKWARD or FORWARD — so a 29 February birthday can land on 1 March in non-leap years instead of vanishing (RFC 7529).

Staff insight: The standard's default for a "monthly on the 31st" rule is to skip the seven months without a 31st, and its default for a recurring 02:30 meeting on spring-forward day is to drop that instance, while a single event at 02:30 is moved to 03:30. Users rarely expect either. Name the policy you implement, implement it identically on every client, and say so in the interview — it is the fastest way to show you have actually expanded RRULEs.

IANA Time Zone Database — Rules Change on Days of Notice#

The tz database records "the history of local time for many representative locations" and is updated when political bodies change offsets or DST rules; RFC 6557 (BCP 175, February 2012) makes the TZ Coordinator an IANA Designated Expert who decides on changes and release timing (IANA, RFC 6557). The release history shows the cadence: seven releases in 2022 (2022a–g), four in 2023, two in 2024, three in 2025, and five by the end of September 2026. Release 2022f, published 28 October 2022, recorded that Mexico would stop observing DST after 2022 (except near the US border) and that Chihuahua would move to year-round −06 on 30 October — two days later. Release 2023b (23 March 2023) moved Lebanon's spring-forward from 25/26 March to 20/21 April; 2023c, five days later, reverted it. In 2026, release 2026d recorded that Canada's Northwest Territories would not fall back on 1 November 2026, and 2026e (29 September 2026) that Manitoba stays on −05 (tz NEWS).

Staff insight: A calendar is the system most exposed to these releases, because it holds future local times. Two-day notice means the rebase of future instances, busy intervals and reminders must be an automated pipeline measured in hours, and that server and devices will disagree until OS vendors ship the same data. Design for skew, not for agreement.

CalDAV, iTIP and iMIP — Sync and Scheduling as Standards#

CalDAV (RFC 4791, March 2007) extends WebDAV with a calendar-query REPORT with time-range filtering and optional expansion of recurring events, calendar-multiget, and a free-busy-query REPORT that returns availability without event details (RFC 4791). iTIP (RFC 5546, December 2009) defines the scheduling methods — PUBLISH, REQUEST, REPLY, ADD, CANCEL, REFRESH, COUNTER, DECLINECOUNTER — with UID as the primary key and an organizer-incremented SEQUENCE so a higher sequence obsoletes earlier versions (RFC 5546); iMIP (RFC 6047) carries those messages over email (RFC 6047). CalDAV Scheduling (RFC 6638, June 2012) moves iTIP delivery to the server ("implicit scheduling") with scheduling inbox and outbox collections (RFC 6638), and WebDAV sync (RFC 6578, March 2012) adds sync tokens, with a DAV:valid-sync-token error that sends clients back to a full sync when the server no longer has the history (RFC 6578).

Staff insight: The standards already encode the Staff answers: the organizer owns the event and SEQUENCE orders versions; free/busy is a separate, detail-free query; sync is a cursor that can expire. If you support CalDAV you inherit clients that poll on their own schedule and send whole iCalendar objects back — so your internal model must round-trip iCalendar without losing exceptions.

Google Calendar API — Sync Tokens, Bodiless Pushes, Bounded Free/Busy#

Google's Calendar API documents incremental sync with nextSyncToken, present only on the last page of a full sync; when a token is invalidated — "token expiration or changes in related ACLs" — the server returns 410 and the client should wipe its store and do a full sync (Google sync guide). Push notifications carry no body, only headers such as X-Goog-Resource-State; channels expire and must be replaced by calling watch again; and the guide warns that notifications "are not 100% reliable" and a small percentage are dropped (Google push guide). The free/busy query accepts calendarExpansionMax up to 50 and groupExpansionMax up to 100 and returns only busy ranges per calendar (freebusy.query). Recurring-event instances carry recurringEventId and originalStartTime (Google recurring events guide).

Staff insight: This is the Staff design as an API contract: push is a hint and polling with a cursor is the guarantee; an expired cursor costs a full resync, so cursor retention is a capacity decision; free/busy is bounded per request so one query cannot fan out to an entire company; and instances are identified by their original start. Cite the 410 behavior when asked what happens to offline clients.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Store times in UTC""Chile changes its DST date next week. What happens to a recurring 09:00 meeting there?"Intent vs instant; tz rebase
"Recurring events are expanded into rows""The organizer moves the series 30 minutes later. What happens to the instance someone already moved?"Exceptions keyed by original start; series edits
"Each attendee gets a copy""The organizer edits the title while an attendee declines. Which write wins?"Ownership split; overlay rows; SEQUENCE
"We query attendees' calendars for free/busy""Find a 30-minute slot for 12 people and a room next week. How many reads is that?"Busy-interval index; merge cost
"We schedule a reminder when the event is created""The attendee declines, then the organizer moves it. Does the reminder still fire?"Reminders as a re-derived projection
"Clients sync via the API""An iPhone has been offline for six weeks. What does it fetch when it reconnects?"Change-log retention; token expiry; full resync cost
"Rooms are attendees""Two people book Room 4B for 10:00 at the same moment. Who gets it?"Conditional write on the room's busy set

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The event store holds intent — local times, zones, rules, exceptions — and is the only source of truth. Everything to its right is derived: the busy-interval index (UTC instants for the next 18 months), the reminder timers (next 48 hours), the per-calendar change logs that clients sync from, and outbound iTIP messages. Derivation runs off a change stream, so an edit is one transactional write plus asynchronous projection. The TZ rebase job is the piece most designs omit: it re-derives future instants when the tz database changes. Reminder delivery hands off to Notifications; timer firing follows Job Scheduler.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Time"Store UTC""Past in UTC. Future as local time + IANA zone; derived instants carry a tzdb version; tz releases trigger a rebase."
Recurrence"Insert a row per instance""Store RRULE + exceptions keyed by original start; expand on read; materialize busy intervals 18 months out."
Invites"Copy to each attendee""One organizer-owned event, per-attendee overlay rows; iTIP with SEQUENCE for external attendees."
Free/busy"Query each calendar""Busy-interval index per calendar, merged per query; intervals only, no titles; ≤ 5 s staleness SLO."
Rooms"Rooms are attendees""Rooms are resources with a no-overlap constraint; booking is a conditional insert, not a check-then-write."
Reminders"Schedule at create""Rolling 48 h timer projection re-derived on edit, RSVP and tz change; dedupe key includes SEQUENCE."
Sync"Clients call the API""Per-calendar change log with tokens; push is a hint, poll is the guarantee; 30-day retention, then resync."

Numbers to Bring#

MetricValueWhy It Matters
Users / daily actives (design assumption)500M / 150M; ~60M enterprise seatsSets every rate below
Event writes (create, edit, RSVP)~1B/day; ~12K/s avg, ~60K/s Monday-morning peakWrite path is modest; fan-out behind it is not
Calendar view / agenda reads~2B/day; ~100K/s peakEach read expands rules for a 1–5 week window
Sync "anything changed?" checks~400M connected clients ÷ 15 min backstop ≈ 450K/sThe single largest request class; must be a cheap cursor compare
Free/busy queries~1.5B/day; ~80K/s peak; ~6 calendars × 2 weeks eachWhy an index beats expanding rules per query
RRULE expansion cost~1–5 µs per instance in a compiled engineA year of a daily series is ~2 ms: fine per read, not per free/busy probe
Busy-index horizon18 months materialized; beyond that, expand on readCovers nearly all scheduling; keeps the index ~70 KB per active calendar
Reminder rate~1.2B/day server-side; minute at :50 ≈ 20–30× averageThe top-of-hour herd: most meetings start on :00 or :30
Reminder lateness targetp99 ≤ 30 sUsers notice a 10-minute reminder arriving at 8 minutes
tz database releases2–7 per year (seven in 2022, two in 2024) (tz NEWS)Rebase must be routine, not an incident
Shortest tz notice seen~2 days (Chihuahua, tz 2022f) (tz NEWS)Rebase SLO must be hours, end to end
Change-log retention for sync30 days (design choice)Longer-offline clients pay a full resync; shorter retention means resync storms
Free/busy staleness SLO≤ 5 s from write to indexBeyond that, people double-book rooms they "saw" free
Google free/busy request boundscalendarExpansionMax ≤ 50, groupExpansionMax ≤ 100 (freebusy.query)Bound fan-out per query; groups are not a free lunch

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on the month-view UI, an events table and a REST API, then meet the first hard question — "what happens to this weekly meeting when the organizer's country drops daylight saving?" — with no answer, because they already threw away the zone. Compress the CRUD to ~8 minutes and spend the rest on the time model, recurrence, the invite ownership split, free/busy and reminders.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Users create single and recurring events, invite people inside and outside the company, and those people RSVP. Anyone scheduling a meeting can see when attendees and rooms are busy without seeing what they're doing. Reminders fire before events by push, email or on the device. Phones, web and third-party clients stay in sync, including offline. Hosts can publish a booking page where others pick a free slot."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: time correctness — an event shows the same local time on every device, survives DST, and survives governments changing the rules; that decides the storage model. Two: recurrence — series are infinite, so range queries must work without infinite rows. Three: free/busy for a meeting with a dozen people and a room must come back in under 200 ms and be fresh within about 5 seconds, or people double-book. Four: reminders must fire within 30 seconds of the right moment and must follow every edit. Scale: 500 million users, 150 million daily, 60 million enterprise seats, about a billion event writes a day, and roughly 400 million connected clients syncing."

Then name the underspecified parts:

"A few things I'd confirm: is this consumer, enterprise, or both? Do we need to interoperate with Exchange and CalDAV clients, or only our own apps? Are rooms and other resources in scope? Are booking pages with payments in scope? I'll assume both consumer and enterprise on one core, CalDAV and email interop, rooms in scope, and booking pages as a layer on top."

🎯 Staff Move: Saying "for future events, the source of truth is the local time and the zone; every UTC instant is a cache" in the first three minutes tells the interviewer you have already decided the hardest correctness question. Everything you draw next can be judged against it.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Calendar: calendar_id, owner (user, group, room, resource), default_tz (IANA ID), acl[] (free/busy only, read details, write, delegate)
  • Event (series or single): event_id, ical_uid, calendar_id (organizer's), start_local, end_local or duration, tzid, all_day, rrule, rdates[], exdates[], sequence, title, location, visibility, status
  • Exception (instance override): event_id, original_start (RECURRENCE-ID), overridden fields only, cancelled
  • Attendance (overlay): event_id, attendee (user, room or external email), partstat (needs-action, accepted, declined, tentative), role, personal_reminders[], hidden, last_seen_sequence
  • Busy interval (derived): calendar_id, start_utc, end_utc, event_id, original_start, transparency, tzdb_version
  • Reminder timer (derived): fire_at_utc, user_id, event_id, original_start, offset, channel, sequence
  • Change log entry: calendar_id, seq (monotonic per calendar), event_id, op

API:

POST  /v1/calendars/{cal}/events            { start_local, tzid, duration, rrule?, attendees[] }
PATCH /v1/events/{id}?scope=all|this|following&original_start=...   (If-Match: sequence)
POST  /v1/events/{id}/rsvp                  { partstat, original_start? }     (overlay write only)
GET   /v1/calendars/{cal}/events?from=...&to=...&expand=true          (range read, expanded)
POST  /v1/freebusy   { calendars[≤50], from, to }  → { cal: [ {start_utc, end_utc}, ... ] }
POST  /v1/rooms/{room}/book   { event_id, original_start, idempotency_key }
GET   /v1/calendars/{cal}/changes?sync_token=...   → { changes[], next_sync_token } | 410

The scope parameter on PATCH is the most important field on this page: "this instance", "this and following" and "all" are three different writes with three different blast radii.

🎯 Staff Move: "An RSVP is a write to the attendee's overlay row, never to the event. That one choice means the organizer editing the title and 300 attendees answering at the same time never contend on the same row — and it's also exactly how iTIP separates REQUEST from REPLY."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one edit in 90 seconds:

  1. Organizer in London creates "Design review, Tuesdays 09:30, Europe/London, FREQ=WEEKLY;BYDAY=TU" with 12 attendees and Room 4B. The event service writes one series row, 13 overlay rows, and one change-log entry per affected calendar — in one transaction on the organizer's shard, with overlays for other shards written through the change stream.
  2. The room is booked by a conditional insert of the first 18 months of instances into the room's busy set; any overlap rejects that instance and the organizer sees which dates conflict.
  3. Projectors expand the rule for 18 months, convert each instance to UTC with the current tzdb, and write busy intervals for every attendee who hasn't declined.
  4. The reminder projector writes timers for instances in the next 48 hours; a sweeper extends the horizon hourly.
  5. Attendees' devices get a push hint, call changes?sync_token=…, receive the series and their overlay, expand locally, and schedule local alarms.
  6. External attendees get an iMIP REQUEST with UID and SEQUENCE 0; their REPLY comes back through the iMIP gateway and updates only their overlay row.
  7. A New York attendee sees 04:30 in winter and 05:30 for the two or three weeks in March (and the one week around early November) when the US and UK are out of step on DST — because London is the anchor, which is what the organizer meant.

🎯 Staff Move: Say out loud: "Notice that only one box holds truth. The busy index, the reminders, the change logs and the outbound invites are all projections of the event store. If any of them is wrong — after a bug, a tz release, or a missed message — I can rebuild it from intent. That's the property I'll protect for the rest of the interview." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes a calendar hard is that most of what it stores is in the future, and the future can change underneath it. I'd like to go deep on four things: how we represent time so DST and tz rule changes don't corrupt events; how recurrence and exceptions work when someone edits 'this and following'; how free/busy and room booking stay correct across a dozen people; and how reminders and device sync follow every change. Where would you like to start?"

If no preference: start with the time model. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: The time model (7–8 min)

"Three kinds of time, stored three ways. A future timed event is a local start, a duration, and an IANA zone ID — the organizer's zone unless they pick another. An all-day event is a date or date range, half-open, with no zone; Christmas is the 25th everywhere. A floating event — 'medication at 08:00' — has a local time and no zone, belongs only to its owner, and never enters a shared free/busy index. Once an instance is in the past, I freeze its UTC instant: an audit of 'when did this meeting happen' must not change because a tz release corrected history."

Then DST, precisely: "Two edge cases. Spring-forward: 02:30 on the transition day doesn't exist in New York. RFC 5545 says a recurring instance generated there is dropped; a single event at 02:30 is read with the pre-gap offset, landing at 03:30. I'd shift recurring instances forward too — users don't expect a meeting to vanish — and I'd document that as our policy and test it in every client. Fall-back: 01:30 happens twice; we take the first occurrence, as the RFC does."

Quantify: "A tz release changes a handful of zones. If a changed zone holds 5 million users with ~40 active series each, that's 200 million series to re-expand, but only future instances in the 18-month busy horizon — roughly 15 billion interval rows at worst, realistically far fewer because only instances after the effective date move. At 1 million rows a second across the projector fleet, about four hours. That's the SLO I'd publish: tz release to consistent index in under 6 hours."

Who pays: "Users in the affected zone pay a window where the server has new rules and their phone doesn't — or the reverse. I'd render server-computed times in our apps for affected zones until device tz data catches up, and show an in-app notice. Platform pays for the rebase pipeline and the version-skew dashboard."


Deep dive 2: Recurrence, exceptions and series edits (7–8 min)

"A series is a row with an RRULE plus RDATE and EXDATE lists. An exception is a row keyed by (event_id, original_start), holding only the overridden fields. Range reads expand the rule into the window, drop EXDATEs, add RDATEs, and overlay exceptions by original start. Three edit scopes: 'this instance' writes or updates one exception; 'all' rewrites the series and bumps SEQUENCE, keeping exceptions whose original start still matches the rule; 'this and following' splits the series — the old one gets UNTIL just before the split point, the new one starts there with a new ID linked by split_from, and exceptions after the split are re-homed to the new series."

Quantify: "Expansion is 1–5 µs per instance. A week view with 40 series is 40 instances — microseconds. The expensive case is a daily series with no end opened in a year view: 365 instances, about 1 ms. I cap expansion per request at 5,000 instances and paginate beyond it."

Who pays: "When the organizer changes a series' time, exceptions keyed to old original starts would be orphaned. I keep them if their original start is still generated by the new rule, and otherwise surface 'n instances had custom changes that were discarded' to the organizer before they confirm. The organizer pays a confirmation dialog; attendees don't pay with ghost meetings."


Deep dive 3: Free/busy and rooms (6–7 min)

"Free/busy reads a per-calendar index of busy intervals in UTC — instance-level, already expanded, already filtered for declined RSVPs and transparent events. A query for 12 people and 3 candidate rooms over two weeks is 15 range scans on 15 shards, merged into free windows in the requester's zone. Each scan returns maybe 40 intervals. The response contains intervals, never titles, unless the requester has delegate access."

"Rooms are different: a room is a resource with a hard invariant — no overlapping accepted bookings. Booking a room is a conditional insert into the room's interval set, enforced by the store — in PostgreSQL an exclusion constraint on (room_id, tstzrange) — not a free/busy check followed by a write. That's the Hotel & Home Booking answer applied to 30-minute slots."

Quantify: "80K free/busy queries a second at ~6 calendars each is ~500K range scans a second, each a few KB from a hot partition. The index is ~70 KB per active calendar for 18 months — about 35 TB across 500M calendars before replication, with the active working set far smaller."


Deep dive 4: Reminders and sync (5–6 min)

"Reminders are a projection: for every instance in the next 48 hours, for every attendee who hasn't declined, for every reminder they've configured, one timer keyed by (user, event, original_start, offset). Any edit, RSVP, or tz rebase re-derives timers for that event; the timer carries the event SEQUENCE so a stale timer that fires after an edit is dropped at delivery. Firing is a Job Scheduler problem; delivery is a Notifications problem. Devices also schedule local alarms from synced data, so a phone in airplane mode still rings; the server suppresses push for devices that confirm local scheduling."

"Sync is a per-calendar change log with a monotonic sequence. A client holds a token per calendar; 'anything changed?' is a compare against the latest sequence in a cache — 450K/s, almost all answered 'no' from memory. Push is a hint; the 15-minute poll is the guarantee. Tokens older than 30 days get a 410 and a full resync — paced, because a bug that invalidates tokens for 10 million clients is a self-inflicted outage."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: intent is the truth, instants are a cache. Future events are local time plus IANA zone plus rule; past events freeze in UTC. Series are rules with exceptions keyed by original start. One organizer-owned event with per-attendee overlays, iTIP at the boundary. Free/busy, reminders and sync logs are projections that a tz release, an edit or a bug can rebuild. And rooms get a hard no-overlap constraint, not a check-then-write."

The evolution closer:

"What I'd build later: a shared recurrence engine with a conformance suite across server, web and mobile; a tz version-skew dashboard; smart scheduling that suggests times using the same busy index; and booking pages with payment holds on the reservation path. What I'd not build: our own tz database or a custom invite protocol — iCalendar, iTIP and CalDAV already exist and every client speaks them."

🎯 Staff Move: End on the correctness failure you designed out and who owns it. Senior candidates end with "and we shard by user." Staff candidates end with "and cal.tzdb.rebase_lag_hours, cal.freebusy.staleness_s and cal.reminder.lateness_p99 are on the calendar platform's dashboard, because those three numbers tell us whether people are showing up at the right time."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
UI tourDesigns day/week/month views and drag-and-dropOne sentence: "Clients render ranges from the range-read API"
UTC by reflex"All timestamps UTC" in Phase 1"UTC for the past; local + zone for the future" in Phase 1
Row-per-instanceMaterializes two years of every series"Store the rule; materialize only busy intervals, 18 months"
Invites as copiesFans every edit out to N calendars"One event, overlays per attendee" in Phase 2
Free/busy as a joinExpands rules for 12 calendars per queryNames the busy-interval index and its staleness SLO
Reminders as an afterthought"And a cron job sends reminders" at minute 44Reminders as a projection with a dedupe key and a lateness SLO

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

A calendar is the cleanest example of a system whose data means something different tomorrow than it does today. Almost every other system stores facts: a payment happened, a message was sent, a file has these bytes. A calendar stores promises about the future in civil time, and civil time is defined by legislatures, not by physics. That makes the obvious engineering default — normalize everything to UTC — a correctness bug for exactly the data the product exists to hold.

It also has a deceptive happy path. A Senior engineer can build a calendar that works perfectly in a demo: a table, a REST API, a week view, invites. The design is judged entirely by days nobody demos: the weekend the US switches to DST and Europe hasn't, the Thursday a government announces it is dropping DST on Sunday, the morning the organizer drags a series with 40 exceptions to a new time, and the 08:50 when a third of the company's meetings fire reminders in the same second. Staff candidates design for those days before drawing the first box — and they notice that most of the system is derived views of a small amount of intent, which is what makes it repairable.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"An organizer in Phoenix creates a weekly 09:00 meeting with attendees in Denver and London. Arizona doesn't observe DST; Denver does; the UK does, on different dates. What time does each attendee see in the second week of March, the first week of April, and the first week of November — and what do you store so that a future change to any of those rules is handled correctly?"

A candidate who answers with the anchor (the meeting is 09:00 America/Phoenix, which is always UTC−7, so the instant never moves), the Denver attendee (09:00 MST in winter, 10:00 MDT once the US springs forward on the second Sunday of March, and back to 09:00 after the first Sunday of November), the London attendee (16:00 GMT in winter, 17:00 BST from the UK's spring-forward in late March), and the storage (local time, America/Phoenix, RRULE — and derived UTC instants tagged with the tzdb version, recomputed if Arizona's rules ever change) has built a calendar. A candidate who says "we store UTC and convert on display" has built a database of timestamps — and cannot tell you what the meeting was supposed to be.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Personal and shared calendars (consumer). The population is individuals and families: 500 million users, several devices each, intermittent connectivity, and a long tail of third-party clients. Most events are created on phones; many are recurring (birthdays, classes, routines); many are all-day or floating. The threats are correctness drift between devices, reminders that fire wrong or twice, and sync that loses an offline edit. The design centers on the intent-based time model, expand-on-read recurrence (so phones can expand locally too), a change-log sync protocol with tokens, and device-local alarms backed by server reminders. Correctness bar: every client shows the same local time for every instance; reminder lateness p99 ≤ 30 s; no lost offline edits.

Enterprise scheduling (free/busy, rooms, delegation). The population is companies: 60 million seats, executives with assistants who manage their calendars, thousands of rooms per campus, and meetings that cross time zones and company boundaries. The core query is "when are these 12 people and one room all free in the next two weeks?" The threats are double-booked rooms, free/busy that lags edits, private event details leaking through availability, and Exchange or CalDAV interop that silently drops updates. The design centers on the busy-interval index, organizer-owned events with attendee overlays, rooms as resources with a hard no-overlap invariant, an ACL model with "free/busy only" as a first-class permission, and iTIP for external attendees. Correctness bar: free/busy fresh within 5 s, zero double-booked rooms, no detail leaks.

Booking pages and appointment scheduling. The population is hosts (consultants, clinics, recruiters, sales teams) and the strangers who book them. A host publishes availability rules — working hours in their zone, buffers, minimum notice, maximum bookings per day — and a booker picks a slot shown in their zone. Traffic is bursty: a link shared in a newsletter produces thousands of availability reads in minutes for one host — a hot key. The design centers on computing availability from the same busy index, then committing a booking as an idempotent conditional insert that also blocks the host's calendar, with optional short holds while a payment completes. Correctness bar: zero double bookings, slot times correct across zones, booking confirmation in under 2 s.

🎯 Staff Move: "These share the time model, recurrence and the busy index, and almost nothing else. Consumer calendars need offline sync and local alarms; enterprise needs privacy-aware free/busy and room invariants; booking pages need a reservation-style write path under bursty load. I'll build one core and give each intent its own write and read policies on top."

2.2 When NOT to Build a Calendar Service#

SituationWhat to Do InsteadWhy
Your product needs scheduling but calendars aren't the productIntegrate with users' existing calendars via their providers' APIs or CalDAVUsers won't move their calendar to you; you need their free/busy, not their events
Appointment booking for a small businessBuy a booking product that writes to the host's calendarRecurrence, tz and interop are years of edge cases for a feature, not a business
System jobs that run "every night at 02:00"A job scheduler — see Job SchedulerMachines don't RSVP; a calendar's DST policy (shift, don't drop) is wrong for jobs that must not run twice
Hotel nights, rentals, inventoryA reservation system with date-based inventory — see Hotel & Home BookingNights are dates with counts and holds; calendar events have no inventory model
Shift rosters with labour rulesA workforce-management system that exports to calendarsOvertime, rest periods and union rules are constraints a calendar doesn't model
Reminders with no time-of-day semantics ("ping me in 2 hours")A delayed message on a queueRelative delays are UTC durations; no zone, no recurrence, no calendar
Public event listings (concerts, sports fixtures)Publish an ICS feed or an "add to calendar" linkReaders subscribe; you never need their availability or RSVP

And within the design, some things you should not build even when you own the service:

  • Don't invent a recurrence format. Use RFC 5545 RRULE and its exception model; every client and every import/export already speaks it.
  • Don't maintain your own tz data. Ingest IANA tz releases; your job is ingestion speed and skew measurement, not geopolitics.
  • Don't put titles in the free/busy index. Keep it intervals only, so a bug in an ACL check cannot leak "Layoff planning" to the whole company.
  • Don't treat rooms as people. A person can be double-booked and choose; a room cannot. Different invariant, different write path.
  • Don't fire reminders from the event row. A timer created at event creation goes stale on the first edit; derive timers from current intent.

The Staff signal is knowing that a calendar is the right tool for human commitments in civil time and the wrong tool for machine schedules, inventory and relative delays — and that most calendar bugs come from mixing those three in one model. See Buy or Build: The Total-Cost Test and Drill 7.

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Which zone anchors a meetingDecides what moves when DST differs between attendees"The organizer's zone by default; the organizer can pin another; attendees see it converted."
Interop scopeExchange/CalDAV/iMIP adds round-trip and ordering constraints"CalDAV and iMIP first-class; ICS subscribe read-only; Exchange via its API."
Free/busy freshnessDecides synchronous vs projected index"≤ 5 s staleness; rooms checked synchronously at booking."
Rooms and resourcesHard invariant vs soft availability"Rooms get a no-overlap constraint; people don't."
Max attendeesFan-out of edits and reminders"1,000 interactive guests; above that it's a broadcast event without per-guest RSVP tracking."
Offline durationChange-log retention and resync cost"30 days of change log; older clients do a paced full resync."
Privacy modelWho sees titles vs busy blocks"Free/busy-only by default inside a company; delegates see details; nothing outside."

2.4 Precise Terminology#

TermMeaningCommon Confusion
InstantA point on the UTC timelineAssumed to capture what the user meant
Wall-clock (local) timeClock reading in a place, e.g. 09:30Treated as unique; it repeats in fall-back and skips in spring-forward
IANA zone IDEurope/London, America/Chicago — a named rule historyConfused with abbreviations (CST is ambiguous) or fixed offsets (−06:00 has no DST)
UTC offsetDifference from UTC at an instantAssumed constant for a zone
DST gapLocal times skipped at spring-forward (e.g. 02:00–02:59)Events there "don't exist" and need a policy
DST overlapLocal times repeated at fall-back (e.g. 01:00–01:59 twice)Ambiguous; RFC 5545 picks the first
Floating timeLocal time with no zoneShared by mistake; means a different instant for each viewer
All-day eventA date or date range, not an instantStored as midnight UTC, appears on the wrong day west of Greenwich
RRULERecurrence rule (RFC 5545)Expanded once and thrown away
Original start / RECURRENCE-IDThe start an instance would have had from the ruleConfused with its current (moved) start
ExceptionOverride for one instance, keyed by original startOrphaned when the rule changes
Series split"This and following": old series ends, new one beginsImplemented as editing every future row
SEQUENCEOrganizer's revision number for an event (iTIP)Incremented by attendees, or not at all
PARTSTATAn attendee's participation statusStored on the event, contending with organizer edits
Free/busyBusy intervals without detailsConfused with read access to the calendar
TransparencyWhether an event blocks time (OPAQUE) or not (TRANSPARENT)All-day "working from home" blocks everyone's whole day

3. Where the Design Splits#

Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.

3.1 Fault Line 1: Recurrence — Store the Rule vs Materialize Instances#

The tension: A rule is compact, editable in one write, and infinite. Rows are trivially indexable and queryable by range. Every calendar has to answer range questions over infinite rules; the only choice is where expansion happens and how much of it is kept.

StrategyWhat WorksWhat BreaksWho Pays
Materialize every instance (N years)Range queries are plain index scans; per-instance edits are row updatesSeries "end" at the horizon; a time change rewrites N rows; exceptions become indistinguishable from generated rows; 40 attendees × daily × 2 years ≈ 29K rows per clickOrganizers (lost exceptions); storage; every edit's write amplification
Store rule, expand on every readOne row per series; edits are one write; infinite series are freeFree/busy for 12 people expands 12 rule sets per query; every client needs an identical expanderFree/busy latency; client teams (N expanders)
Store rule + exceptions; expand on read; materialize busy intervals for a bounded horizon (Staff default)Edits are one write; range reads expand cheaply; free/busy reads an index; the index is rebuildableTwo representations to keep consistent; horizon must roll forward; tz changes must rebasePlatform owns the projector and horizon sweeper
Client-only expansion (CalDAV-style "give me the master")Server is simpleServer can't answer free/busy or fire reminders without expanding anywayEvery client; server features you can't build
Diagram: 3.1 Fault Line 1: Recurrence — Store the Rule vs Materialize Instances

The subtle part — moved exceptions. An exception can move an instance into a window whose rule-generated instances don't include it (Tuesday's meeting moved to the following Monday), or out of the window its original start falls in. A correct range read therefore also queries exceptions whose current start overlaps the window, not only those whose original start does. Index exceptions by both.

Series edits, precisely:

Edit ScopeWriteExceptionsSEQUENCE
This instanceUpsert exception at original_startOnly this oneIncremented for the event (iTIP sends a REQUEST with RECURRENCE-ID)
All instancesUpdate series rowKept if their original start is still generated; otherwise listed to the organizer and dropped on confirmIncremented
This and followingSet UNTIL on old series to just before split; create new series with split_from; move exceptions after split to the new seriesRe-homedNew series starts at 0; old series incremented

RFC 5545 also allows RECURRENCE-ID;RANGE=THISANDFUTURE for "this and following". Inside the system I prefer the split, because it keeps every series' history immutable before the split point and makes each series a simple rule; at the CalDAV and iMIP boundary I translate either form.

The Staff default: store the rule, RDATE/EXDATE and exceptions keyed by original start; expand on read in a single shared engine; materialize busy intervals and reminders for bounded horizons (18 months and 48 hours) as projections.

When to deviate:

  • Room and resource calendars: materialize instances into the room's booking table for the booking horizon, because the no-overlap constraint must be enforced per instance at write time (see 3.4).
  • Very long horizons in reporting ("how many hours of meetings will we have next year?"): expand in a batch job, not the online index.
  • Series with thousands of exceptions (a daily standup with years of moves): split the series automatically at a year boundary to bound per-read work.

3.2 Fault Line 2: Time Representation — UTC Instant vs Wall-Clock + Zone#

The tension: UTC instants compare, sort and index trivially and are the right default almost everywhere in backend engineering. But a future calendar event is a promise in civil time, and civil time's mapping to UTC is decided by governments and can change after you store it. Storing only the instant throws away the one thing you need to apply that change.

StrategyWhat WorksWhat BreaksWho Pays
UTC instant onlySimple; index-friendlytz rule change → every future instance wrong, unrecoverably; DST drift for recurring series computed with old rulesUsers in the affected zone; support; a migration that can't recover intent
Local time + fixed UTC offset (09:00-06:00)Looks preciseNo DST: a winter offset applied in summer is wrong by an hourEvery user with a recurring event across a DST change
Local time + IANA zone, instants derived on readSurvives rule changesNo index: every free/busy query convertsFree/busy latency
Local time + IANA zone as truth; UTC instants derived into an index tagged with tzdb version; past frozen (Staff default)Correct after rule changes; fast range queries; auditable pastA rebase pipeline; version skew between server and devicesPlatform owns rebase and skew metrics
Diagram: 3.2 Fault Line 2: Time Representation — UTC Instant vs Wall-Clock + Zone

DST, precisely. In the US, clocks spring forward from 02:00 to 03:00 on the second Sunday of March and fall back from 02:00 to 01:00 on the first Sunday of November; the EU and UK switch on the last Sundays of March and October. That gives a two-to-three-week window in March and a one-week window around the end of October when transatlantic meetings shift by an hour for one side only — correct behavior, and the most common "the calendar is broken" support ticket.

CaseExampleRFC 5545 RuleOur Policy
Single event in a gap02:30 New York on spring-forward dayUse the pre-gap offset → 03:30 EDTSame, and warn at creation
Recurring instance in a gapDaily 02:30 seriesInstance is ignored and not countedShift to 03:30 (users don't expect a missing instance); documented and tested everywhere
Time in an overlap01:30 on fall-back dayFirst occurrence (EDT)Same; offer "second occurrence" only via explicit UTC
Invalid dateMonthly on the 31st; yearly on 29 FebIgnoredKeep RFC default for BYMONTHDAY=31; offer "last day of month" (BYMONTHDAY=-1) in the UI; birthdays use RSCALE SKIP=FORWARD where clients support it

tz data changes. Every derived row records tzdb_version. When a release arrives, the rebase job diffs old and new compiled rules per zone, finds zones whose offsets differ for any instant after now, and re-derives future instances in those zones only. Past instances — anything that ended before the release was ingested — keep their frozen UTC values. See Section 4.1 for the timeline when notice is two days.

Who signs off: product owns the gap/overlap policy and its wording in the UI; the calendar platform owns the rebase SLO and the tzdb ingestion; client teams own shipping the same policy. That's a written table, not a per-client judgment.

🎯 Staff Move: "I'm storing the meeting the way the organizer said it — 09:30, Europe/London — and treating every UTC instant as a cache tagged with the tz data version that produced it. The past I freeze. That one decision is why a government's two-day notice is a pipeline run for us and a data migration for everyone else."

3.3 Fault Line 3: Invitations — Shared Event vs Per-Attendee Copies#

The tension: Giving every attendee their own copy makes reads local and works naturally across systems — it is how email-based iTIP works between companies. But inside one system it turns every organizer edit into N writes that can partially fail, and makes the question "which version is true?" a distributed-consistency problem.

StrategyWhat WorksWhat BreaksWho Pays
Full copy per attendee (fan-out on write)Each calendar read is local; attendees can annotate freelyEdit = N writes; partial failure leaves stale copies; 1,000-guest event = 1,000 writes per typo fixAttendees who see the old time; on-call reconciling copies
One shared event, no per-attendee stateOne write per editRSVPs, personal reminders and "hide this" all contend on one rowOrganizer (lock contention), attendees (no personal settings)
One organizer-owned event + per-attendee overlay rows (Staff default)Edit = one write; RSVP = overlay write; no contention between them; one truthReads join event + overlay; cross-shard projections for attendees' indexes and logsPlatform owns projection; reads pay a join (cheap, same request)
iTIP copies at the boundary (external attendees)Interoperable with every calendar systemOrdering, duplicates and lost REPLYs are normalIntegration team; SEQUENCE handling
Diagram: 3.3 Fault Line 3: Invitations — Shared Event vs Per-Attendee Copies

Ownership, precisely: the organizer's calendar owns start, end, tzid, rrule, title, location, attendee list and SEQUENCE. Each attendee owns partstat, personal reminders, colour and visibility of their instance. An attendee "moving" a meeting in their own calendar is either a COUNTER proposal to the organizer (iTIP) or a private copy — never a write to the shared event.

Large events. At 1,000+ guests, per-guest RSVP tracking and per-edit fan-out of notifications stop being useful and start being an attack surface. Above that threshold the event becomes a broadcast event: attendees subscribe, RSVP counts are aggregated, edits produce one change-log entry on the event's own feed rather than one per attendee calendar, and the organizer's row stops being a hot key for replies (see 4.3).

The Staff default: one event, organizer-owned, with per-attendee overlays inside the system; iTIP (REQUEST/REPLY/CANCEL with UID + SEQUENCE) at every boundary; broadcast mode above 1,000 guests.

When to deviate: cross-region tenants where an attendee's data must reside in another jurisdiction — store a replica of the event in the attendee's region and treat it as an iTIP-style copy with SEQUENCE ordering, accepting seconds of lag.

3.4 Fault Line 4: Free/Busy — Compute on Read vs Precomputed Busy Index#

The tension: Computing availability from events at query time is always consistent with the source of truth, but it expands every attendee's recurring series for every probe. A precomputed busy index answers in one scan per calendar, but it is a projection that can lag writes, miss a tz rebase, or disagree with the event store after a bug.

StrategyWhat WorksWhat BreaksWho Pays
Expand events per queryAlways exact12 calendars × 40 series × 2-week expansion per query; scheduling assistants issue a query per keystrokeFree/busy latency; event-store read load
Busy-interval index, synchronous update in the write transactionFresh immediatelyEvery edit writes N attendees' index partitions on other shards; distributed transaction per editWrite latency and availability
Busy-interval index via async projection, ≤ 5 s lag (Staff default)Fast reads; writes stay single-shardLag window; needs reconciliationUsers in the lag window; platform owns staleness SLO and reconciler
Index + synchronous check for hard resourcesRooms and booking pages correct at commitTwo paths to maintainBooking service

Rooms are the exception that proves the rule. A person double-booked is an inconvenience they resolve; a room double-booked is two meetings in the corridor. Room booking therefore does not trust the projected index. It is a conditional insert into the room's own instance table, with the invariant enforced by the store:

CREATE TABLE room_bookings (
  room_id        bigint,
  during         tstzrange,             -- UTC, derived from local + zone
  event_id       bigint,
  original_start timestamptz,
  EXCLUDE USING gist (room_id WITH =, during WITH &&)
);

A recurring booking inserts each instance in the 18-month horizon in one transaction; conflicts come back per instance so the organizer sees "Room 4B is taken on 3 of 78 dates". The horizon sweeper extends room bookings as it extends busy intervals; a tz rebase of the room's zone re-derives them inside the same constraint. This is the Hotel & Home Booking inventory invariant applied to time ranges, and the same contention shape when 40 people want the one large room at 10:00 Monday.

Privacy. The index stores intervals, event_id and transparency — no title, location or attendees. Details are fetched only through the event service's ACL check. Free/busy permission is separate from read permission, and external requesters get coarser data (merged blocks, no event IDs).

The Staff default: async-projected busy index with a 5-second staleness SLO for people; synchronous conditional insert for rooms and booking pages; nightly reconciler comparing a sample of calendars' index against fresh expansion and reporting cal.freebusy.divergence.

When to deviate: small tenants (< 10K seats) can compute on read with a per-calendar cache and skip the index entirely; the index earns its complexity at scheduling-assistant query rates.

3.5 Fault Line 5: Reminders — Schedule at Write vs Rolling-Horizon Projection#

The tension: The simplest reminder is a timer created when the event is created. It goes wrong on the first edit, the first decline, the first series change, and the first tz release — and for an infinite series it can't be created at all. A projection that re-derives timers from current intent is always right but costs a re-derivation on every change.

StrategyWhat WorksWhat BreaksWho Pays
Timer created at event creationSimpleStale after edits; can't represent infinite series; reminders for declined eventsUsers (wrong-time and ghost reminders)
Cron scans events every minuteAlways currentExpands every user's rules every minute; 150M calendarsInfra cost; minute-boundary herd
Rolling 48 h projection, re-derived on change (Staff default)Correct after edits, RSVPs and tz rebases; bounded timer countRe-derivation fan-out on big series edits; sweeper must extend horizonPlatform owns projector and sweeper
Device-local alarms onlyWorks offline; zero server costDevices that haven't synced fire stale reminders; email and wearables need the serverUsers with stale devices
Diagram: 3.5 Fault Line 5: Reminders — Schedule at Write vs Rolling-Horizon Projection

Dedupe and staleness. The timer key is (user_id, event_id, original_start, offset_minutes, channel); the timer carries the event SEQUENCE it was derived from. At fire time the worker compares against the event's current SEQUENCE: older means a re-derivation is in flight or missed, so it re-reads intent and either fires with fresh data or drops. That makes the timer store safe to be slightly behind — the check at delivery is the correctness backstop.

The top-of-hour herd. Most meetings start on :00 or :30, and most users keep the default reminder offset, so the minute at :50 carries 20–30× the average reminder rate. Jitter is not acceptable — a reminder that arrives two minutes late is a bug users notice. The answer is pre-staging: the timer service loads the next minute's timers into memory 60 s early, partitions them by user hash across workers, and hands batches to the notification system with a deliver_at; the notification system holds them until the exact second. Capacity is planned for the :50 peak, not the daily average. Timer firing mechanics are in Job Scheduler; per-channel delivery, collapse and quiet hours are in Notifications.

Device vs server. Phones schedule local alarms from synced data for the next 7 days and report "scheduled through sequence N" for each event; the server suppresses push for that device when the device's SEQUENCE matches. Email and wearables always use the server path.

The Staff default: rolling 48-hour timer projection re-derived from intent on every change, SEQUENCE check at delivery, pre-staged minute batches for the :50 herd, device-local alarms with server suppression.

When to deviate: all-day events and floating reminders (09:00 "on the day") are evaluated in the user's current zone, which the device knows better than the server — let the device own those and the server send only an email fallback.


4. When It Breaks#

4.1 Short-Notice Time Zone Change — Two Days to Move Every Future Meeting#

t=-2d 10:00: Government announces the region drops DST this weekend. tz maintainers
             publish a new release the same day (the 2022f shape: two days' notice).
t=-2d 14:00: Without a pipeline: nothing happens until an engineer notices.
             Server still on old tzdb. Busy index, reminders, room bookings
             assume the old rules for 1.2M users in the zone.
t=-1d:       Some phones receive an OS update with new tz data; most don't.
t=0 Sun:     Clocks don't change locally. Server and stale phones think they did.
t=+1d 08:50: Reminders fire an hour wrong for recurring meetings in the zone.
             Free/busy shows rooms free that are booked. Cross-zone attendees
             see times that differ from what locals see.
t=+3d:       Support volume 15× in the region. Engineers find no stored
             intent for series written as UTC rows.

With the Staff design:
t=-2d 11:00: tzdb release detected by the ingestion watcher; compiled; diff shows
             one zone changes offsets from Sunday 02:00 onward.
t=-2d 11:30: Rebase job: 1.2M users, 31M future instances in the 18-month horizon
             re-derived; reminders and room bookings re-derived in the same pass.
t=-2d 13:00: Index tagged with new tzdb version. Clients in the zone are told to
             render server-computed times until their device tzdb matches.
t=+1d:       Meetings, rooms and reminders correct. Skew dashboard tracks devices
             still on the old tzdb; in-app notice for those users.

Why it's hard: notice can be days, the server and every device carry independent copies of tz data, and the change must apply to future instances only. Lebanon in March 2023 is the extreme: a release moved its DST start, and another release five days later reverted it (tz NEWS). The rebase must be idempotent and reversible because the rules can flip back.

Detection: cal.tzdb.release_ingest_lag_hours (upstream release to server ingestion), cal.tzdb.rebase_lag_hours, cal.tzdb.version_skew (share of active clients on an older tzdb per zone), cal.render.server_client_mismatch (client reports of instances whose computed time differs from the server's).

Prevention: automated ingestion within an hour of release; rebase scoped by zone diff; every derived row tagged with its tzdb version; a "render server time" capability in clients; a runbook that includes customer comms.

Owner: calendar platform team (ingestion, rebase); client teams (render fallback); support (comms).

4.2 The :50 Reminder Herd#

t=08:49:00:  Normal reminder rate ~15K/s for the region.
t=08:50:00:  Default 10-minute reminders for every 09:00 meeting become due:
             ~400K timers due within the same second, 25× average.
t=08:50:02:  Timer workers read due timers from the store by time bucket; the
             08:50 bucket is one hot partition. p99 read latency 4 s.
t=08:50:30:  Notification fan-out queue backs up behind the burst.
t=08:52:40:  Reminders for 09:00 meetings still arriving. Users join late;
             complaints "reminders are broken on Mondays".

The Staff design: time buckets are sub-partitioned by user hash (64 sub-buckets per minute) so one minute is not one partition; workers pre-load the next minute's timers 60 s early; batches go to the notification system with deliver_at and are held to the second; capacity is provisioned for the 99th-percentile minute of the week (Monday 08:50 local in the largest regions), not the average. The same hot-bucket problem appears in any time-indexed queue — see Hot Keys and Job Scheduler.

Detection: cal.reminder.lateness_p99 by minute-of-hour, cal.timer.bucket_read_ms, cal.reminder.due_backlog.

Owner: calendar platform (timers); notification platform (delivery capacity at the burst).

4.3 RSVP Storm on a Company-Wide Event#

t=0:       Comms sends a 60,000-person all-hands invite from one organizer
           calendar. Event treated as a normal event: 60K overlay rows, 60K busy
           index writes, 60K change-log entries, 60K push hints.
t=+2min:   18,000 RSVPs in two minutes. Each REPLY also notifies the organizer
           ("Alex accepted"), writing to the organizer's change log.
t=+3min:   Organizer's calendar shard: change-log append is a hot key at 150/s;
           the organizer's own phone receives a sync hint per reply and resyncs
           continuously. Shard p99 for every other calendar on it: 2 s.
t=+10min:  Comms edits the dial-in link. Another 60K-way fan-out + 60K iMIP
           REQUESTs to external partners on the list. Mail gateway throttled.

The Staff design: above 1,000 guests the event becomes a broadcast event: RSVPs aggregate into counters (sharded, like a leaderboard counter), the organizer gets a summary rather than per-reply notifications, attendee calendars reference the event through a subscription rather than per-attendee change-log entries, and edits publish once to the event's feed. External recipients above a threshold get one iMIP message per edit batched over 5 minutes.

Detection: cal.shard.hot_calendar_writes_per_s, cal.event.attendee_count at creation (warn above 1,000), cal.imip.outbound_queue_depth.

Owner: calendar platform; the internal comms tool owner signs off on broadcast mode as the default for distribution-list invites.

4.4 Sync Token Invalidation Storm#

t=0:       A change-log compaction job is deployed with a bug: it truncates logs
           to 2 days instead of 30.
t=+6h:     Clients offline > 2 days present tokens older than the log: 410.
           Each does a full resync: list every event in every calendar it holds.
t=+8h:     Monday morning: 22M clients reconnect after the weekend; 9M get 410.
           Full resync averages 1,800 events × 3 calendars per client.
t=+8h10m:  Range-read tier at 6× normal load; event store read replicas saturated;
           ordinary week-view loads time out. Clients retry full resyncs.

The Staff design: full resync is a planned, paced operation — the server returns 410 with a Retry-After drawn from a token bucket so resyncs are spread over hours; resync fetches a bounded window first (−30 days to +18 months) and backfills history lazily; compaction jobs assert a minimum retention and are canaried on one shard. This mirrors the cursor-reset failures in Cloud File Sync; Google documents the same 410 → full sync contract for its API (Google sync guide), and RFC 6578 defines it for WebDAV sync.

Detection: cal.sync.full_resyncs_per_min, cal.changelog.min_retention_days by shard, ratio of 410s to 200s on the changes endpoint.

Owner: calendar platform (sync); SRE for the canary policy on data-lifecycle jobs.

4.5 Silent Free/Busy Divergence#

A projector bug drops busy intervals for series whose RRULE contains BYSETPOS ("last weekday of the month") after a library upgrade. Nothing errors. Over three weeks, the scheduling assistant suggests slots that collide with monthly reviews; people decline, re-schedule, and blame the assistant. Rooms are unaffected (they use the synchronous path). Detection comes only from the nightly reconciler: it samples 100K calendars, expands from intent, diffs against the index, and reports cal.freebusy.divergence by RRULE feature — the spike is localized to BYSETPOS within one run. Prevention: the recurrence engine's conformance suite runs before any library upgrade, including BYSETPOS, negative BYMONTHDAY, DST gaps and invalid dates; the reconciler's divergence is a paging alert above 0.1%. Owner: calendar platform.

4.6 Room Double-Booked Through a Side Door#

The room booking path enforces the no-overlap constraint, but a facilities tool imports a weekly class schedule straight into room calendars as ordinary events, bypassing the booking service. Forty rooms are double-booked for a term. The fix is the same as in Hotel & Home Booking: every write path to a resource calendar — UI, API, CalDAV, iMIP auto-accept, admin imports — goes through the booking service; resource calendars reject direct event writes. Detection: cal.room.overlap_count from a periodic constraint audit; owner: calendar platform plus the facilities tool owner.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Short-notice tz changecal.tzdb.rebase_lag_hours; cal.tzdb.version_skewEvery future event in the zone, plus cross-zone attendeesAutomated ingest + scoped rebase; server-time rendering fallbackCalendar platform + client teams
Reminder herd at :50cal.reminder.lateness_p99 by minute-of-hourEvery 09:00 meeting in a regionSub-partitioned buckets; pre-staging; deliver_atCalendar + notification platforms
RSVP storm / hot organizercal.shard.hot_calendar_writes_per_sEvery calendar on the organizer's shardBroadcast mode above 1,000 guests; aggregated RSVPsCalendar platform
Sync token stormcal.sync.full_resyncs_per_minRange-read tier; all users' week viewsPaced 410s; windowed resync; retention guardsCalendar platform
Free/busy divergencecal.freebusy.divergence (reconciler)Scheduling suggestions; trustConformance suite; nightly reconcilerCalendar platform
Room side-door writescal.room.overlap_countRooms on a campusSingle write path to resourcesCalendar platform + facilities
Orphaned exceptions after series editcal.series.exceptions_droppedOne series' attendeesConfirm dialog; keep matching exceptionsCalendar platform + client teams
External iTIP out of ordercal.imip.stale_sequence_repliesOne event with external guestsSEQUENCE ordering; REFRESHInterop team

🎯 Staff Insight: The dangerous calendar failures don't page. A meeting that is an hour off returns 200. A busy interval that is missing shows a free slot. A reminder for a declined meeting is delivered successfully. cal.tzdb.version_skew, cal.freebusy.divergence and cal.reminder.stale_sequence_drops are the three metrics that turn silent correctness failures into tickets — make them first-class from day one.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: events, invites, reminders, viewsNames consumer / enterprise / booking intents; commits; states "intent is truth, instants are cache"Asks how many time libraries and tz data copies the company ships, and how skew is measured
TimeUTC everywhere, convert on displayLocal + IANA zone for future, UTC frozen for past, floating and all-day handled; DST gap/overlap policy stated; tz rebase pipelinetz data as a platform dependency with an ingestion SLO and a short-notice runbook across server and clients
RecurrenceRows per instance, or RRULE with ad hoc editsRule + exceptions keyed by original start; three edit scopes; bounded materializationOne recurrence engine with a conformance suite as an org-wide contract
InvitesCopies per attendeeOrganizer-owned event + overlays; iTIP with SEQUENCE at the boundary; broadcast mode for large eventsInterop portfolio: which protocols are first-class, who funds them, and when to retire one
Free/busy & roomsQuery each calendar; rooms as attendeesBusy-interval index with staleness SLO; rooms as conditional inserts with a no-overlap constraintAvailability as a product API others build on, with privacy guarantees written down
Operations"Add monitoring"cal.tzdb.rebase_lag_hours, cal.freebusy.divergence, cal.reminder.lateness_p99, paced resyncsGame days for tz changes and the Monday-morning herd; correctness SLOs reported alongside availability

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Separates intent from instant"For future events the truth is 09:30 Europe/London; the UTC instant is a cache tagged with the tzdb version."
Knows DST edge cases precisely"02:30 on spring-forward day doesn't exist; RFC 5545 drops recurring instances there. I'd shift them and document it."
Keys exceptions correctly"Exceptions are keyed by original start, so moving the series doesn't lose the instance someone already moved."
Splits ownership"An RSVP writes the attendee's overlay, never the event."
Treats derived state as rebuildable"Busy index, reminders and change logs are projections; I can rebuild any of them from intent."
Names the real hot spots"The biggest request class is 'anything changed?' and the biggest spike is reminders at :50."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Store all times in UTC" with no qualificationCannot survive tz rule changes; loses intent for every future event
Materializes recurring events indefinitely or for a fixed N yearsSeries silently end; edits rewrite thousands of rows; exceptions lost
Stores fixed offsets (-05:00) as the zoneWrong by an hour across every DST transition
Rooms booked by "check free/busy, then insert"Race between two bookers; double-booked rooms
Reminders scheduled once at creationFire for moved, declined and cancelled events
No answer for an offline client returning after weeksSync without a cursor-expiry story loses edits or melts the server

5.4 Common False Positives#

  • Fluency in date libraries ≠ a time model. Knowing ZonedDateTime is good; storing its output as UTC for a future meeting is the bug.
  • "We support RRULE" ≠ recurrence design. The questions are exceptions, edit scopes and where expansion happens.
  • Knowing CalDAV exists ≠ sync design. The question is what happens when the token is older than the log.
  • "Rooms are just attendees" sounds elegant ≠ correct. People can be double-booked by choice; rooms cannot.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick intents; "intent is truth, instants are cache"; scale
Entities & API3–5 minEvent, exception by original start, overlay, busy interval, timer, change log; edit scopes
Architecture5–10 min≤ 8 boxes; one source of truth; projections
Time model10–18 minLocal + zone, floating, all-day, frozen past; DST gap/overlap; tz rebase
Recurrence & edits18–25 minExpand on read; exceptions; this/following/all
Free/busy & rooms25–32 minBusy index, staleness SLO, privacy; room constraint
Pivot (interviewer's choice)32–42 minReminders, sync/CalDAV, booking pages, large events, multi-region
Wrap42–45 minIntent vs cache; projections rebuildable; metrics and owners

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"A country drops DST with three days' notice"Time model under changeZone diff → scoped rebase → index, reminders, rooms; device skew fallback
"Find a slot for 15 people and a room"Multi-calendar availabilityBusy index scans, interval merge, working hours in each attendee's zone, room constraint at commit
"Add booking pages like a scheduling link"Availability as a product + contentionRules in host zone; slot display in booker zone; idempotent conditional insert; hot-link caching
"Support Outlook and Apple Calendar users"InteropCalDAV + iMIP; round-trip iCalendar; SEQUENCE; polling clients
"Reminders are late on Monday mornings"Time-bucketed hot spotsSub-partitioned buckets; pre-staging; capacity for the :50 minute
"Go multi-region with data residency"Placement of truthHome region per calendar; cross-region invites as iTIP-style copies

6.3 What to Deliberately Skip#

  • The month-view UI. "Clients render ranges returned by the range-read API."
  • Search over event titles. "A per-user search index fed by the change stream; out of scope."
  • Video-conference links. "An integration that writes a field on the event."
  • Every RRULE part. "FREQ, INTERVAL, BYDAY, BYMONTHDAY, BYSETPOS, COUNT, UNTIL — we use a standard engine."
  • Admin consoles and retention policy UI. "Policy is data the event service enforces."

6.4 Follow-Up Questions to Expect#

  1. "A weekly 09:00 meeting in Chicago has attendees in Berlin. What does Berlin see across March, and why does it change twice?"
  2. "The organizer moves 'this and following' for a series where three later instances were already moved individually. What happens to them?"
  3. "How do you guarantee a room is never double-booked when two people book it at the same time?"
  4. "A user declines one instance of a series. Which systems change, and how fast?"
  5. "A phone has been offline for six weeks. What does it download, and what does that cost us?"
  6. "The tz database ships a fix that changes a zone's past offsets. Do you change past events?"
  7. "How do you show free/busy to someone in another company without leaking event titles?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design Google Calendar."

Staff Answer

"Before drawing — is this consumer calendars, enterprise scheduling with free/busy and rooms, booking pages, or all three? They share a core and differ in write paths. I'll build the shared core and go deepest on enterprise scheduling.

The constraints I'll commit to: future events are stored as local time plus IANA zone plus recurrence rule, and every UTC instant is a derived cache tagged with the tz data version — past events freeze in UTC. Series are rules with exceptions keyed by original start. One organizer-owned event with per-attendee overlays; iTIP at the boundary. Free/busy reads a busy-interval index fresh within 5 seconds; rooms are booked by a conditional insert. Reminders are a rolling 48-hour projection that follows every edit. Scale: 500 million users, 60 million enterprise seats, a billion writes a day, 400 million syncing clients. I'll go: entities → one source of truth plus projections → time model → recurrence and edits → free/busy and rooms → reminders and sync."

Why this is L6:

  • Distinguishes intents and commits
  • States the time model before any boxes
  • Frames free/busy, reminders and sync as rebuildable projections

What L7 adds:

  • Asks how many recurrence engines and tz data copies already exist across server and clients
  • Frames tz data as a company-wide dependency with an ingestion SLO
  • Asks which interop protocols the business has promised enterprise customers
❌ Common L5 Trap

"An events table with start_time and end_time in UTC, a user_id, and an attendees table. Recurring events get expanded into rows for the next year. Reminders are a cron job that queries events starting in the next 10 minutes. Shard by user_id."

Why this misses: Each part works in a demo. The design cannot survive a tz rule change, ends every series after a year, rewrites hundreds of rows per series edit, expands nothing for free/busy, and scans every calendar every minute for reminders.


Drill 2: Core Mechanic — Expanding a Series With Exceptions#

Prompt: "Walk me through rendering next week for a user with a weekly series, one moved instance and one cancelled instance."

Staff Answer

"Load the user's single events whose start overlaps the week, series whose span overlaps it, and exceptions whose original or current start overlaps it. For each series: expand the RRULE in the series' zone between the window bounds, producing local start times; apply the DST policy (gap → shift forward, overlap → first occurrence) and skip invalid dates; convert each to UTC with the current tzdb. Remove instances whose original start is in EXDATE — that's the cancelled one. Overlay exceptions by original start — the moved one takes its new start, which may now fall outside the week, and an exception moved into the week from outside is added because I queried exceptions by current start too. Join the user's overlay for PARTSTAT and personal reminders. Convert to the viewer's zone for display. Cost: tens of instances, microseconds each."

Why this is L6:

  • Expands in the series' zone, not the viewer's
  • Handles exceptions by original start and by current start
  • Applies a stated DST policy during expansion

What L7 adds:

  • Requires every client to use the same engine or pass the same conformance suite
  • Measures cross-client disagreement in production (cal.render.server_client_mismatch)

Drill 3: Make It Concrete — Capacity#

Prompt: "Size it. Storage, reads, the busy index, reminders."

Staff Answer

"Writes: ~1B a day — creates, edits, RSVPs — ~12K/s average, ~60K/s on Monday mornings. Each fans out to projections: on average ~4 attendees, so ~50K busy-index writes a second at average and a few hundred thousand at peak.

Event store: ~1.5 KB per event row, ~120 bytes per overlay; ~200 TB a year of new event data before replication. Shard by calendar_id; an organizer's events and their overlays co-locate on the organizer's shard.

Busy index: an active calendar has ~3,000 instances in 18 months at ~24 bytes each ≈ 70 KB; 500M calendars ≈ 35 TB, most of it cold. Free/busy at 80K queries/s × ~6 calendars = ~500K range scans/s.

Sync: 400M clients; push hints plus a 15-minute backstop poll = ~450K/s 'anything changed?' checks, answered from a cached latest-sequence per calendar — a few hundred bytes in memory per active calendar.

Reminders: ~1.2B/day, ~14K/s average, ~300–400K/s in the minute at :50 in the largest region. 48-hour timer store ≈ 2.4B timers × ~100 bytes ≈ 240 GB."

Why this is L6:

  • Finds the dominant request class (sync checks) and the dominant spike (:50)
  • Sizes derived state separately from truth
  • Ties shard key to ownership (organizer's calendar)

What L7 adds:

  • Prices the device-vs-server split: every reminder a phone fires locally is server capacity not bought for the :50 peak
  • Sets the change-log retention by cost of resyncs vs storage, not by convention

Drill 4: The Dependency Goes Down#

Prompt: "The projector pipeline is down for 40 minutes. What breaks?"

Staff Answer

"Writes still commit — the event store and change log are the truth, and they're on the write path; projections aren't. What goes stale: the busy index (new meetings don't block time), attendee change logs for cross-shard attendees (they don't see new invites), reminders for newly created or moved events within 48 hours, and outbound iMIP. Mitigations in order: free/busy marks calendars with projection lag over 30 s as 'possibly stale' in the UI; room bookings are unaffected because they're synchronous; reminder delivery re-reads intent at fire time, so a moved meeting doesn't fire at the old time — at worst a new meeting's reminder is late, so the timer service also does a just-in-time scan of events created in the last hour when the projector lag alarm is active. When the pipeline resumes it replays from its Kafka offset; projections are idempotent upserts keyed by (calendar, event, original_start)."

Why this is L6:

  • Distinguishes truth from projections and what each failure costs
  • Uses the fire-time SEQUENCE check as a correctness backstop
  • Idempotent replay as the recovery mechanism

What L7 adds:

  • Sets a projection-lag SLO and an error budget, and ties projector deploys to it
  • Runs a quarterly game day that pauses projection on one shard during business hours

Prompt: "A founder posts their booking link to 2 million followers. What happens?"

Staff Answer

"Two load shapes on one calendar. Reads: thousands of availability requests a second for one host. Availability is a pure function of (host rules, busy index, now) that changes only on bookings, so I cache the computed slot list per host per day with a version from the host's change-log sequence — every booking bumps it. The edge serves it; origin computes it once per change. Writes: hundreds of people trying to book the same 20 slots. Each booking is a conditional insert on the host's slot range — same no-overlap constraint as rooms — with an idempotency key from the booker's session, so retries don't double-book and losers get 'slot taken, here are the next three'. I'd add a per-host booking rate limit and a waitlist option. The host's calendar shard sees ~20 successful writes and a few hundred rejected ones; the cache absorbs the rest. This is the Ticket Drops shape at small scale."

Why this is L6:

  • Separates read hot key (cacheable, versioned) from write contention (conditional insert)
  • Idempotency for retries
  • Bounds damage with a per-host limit

What L7 adds:

  • Offers hosts "high-demand mode" (lottery or queue) as a product choice, not an engineering patch
  • Notes the abuse vector: bots booking out a competitor's calendar, and who owns that policy

Drill 6: Multi-Tenant — Free/Busy Across Companies#

Prompt: "Company A wants to see Company B's free/busy for scheduling. How?"

Staff Answer

"Within one provider: a tenant-level sharing policy, set by B's admin, that grants A's users free/busy-only access to B's calendars — merged blocks, no event IDs, no titles, horizon capped at, say, 60 days, rate-limited per requesting tenant. Across providers: the standards path is a CalDAV free-busy-query or a VFREEBUSY request via the scheduling outbox; in practice it's usually provider-to-provider federation configured by both admins. The design property is the same: the busy index stores no details, so a mis-scoped grant leaks availability, never content. Every cross-tenant query is logged for B's admin."

Why this is L6:

  • Free/busy permission separate from read permission
  • Coarser data across trust boundaries
  • Auditable to the data owner

What L7 adds:

  • Treats cross-tenant availability as a contract with legal review and a default-off posture
  • Decides which federation protocols the company commits to supporting for years

Drill 7: Build vs Buy#

Prompt: "We're a recruiting product. Should we build our own calendar?"

Staff Answer

"No. Our users' calendars live in their existing providers, and they won't move. What we need is their free/busy and the ability to write interview events into their calendars. So: integrate with the major providers' APIs plus CalDAV, read free/busy, write events with our own UID namespace, and receive changes via push hints plus sync tokens — treating the 410 'token expired' path as normal. We build what is ours: interview-loop scheduling logic, panel constraints, candidate booking pages. We store a copy of availability with a short TTL and never become the system of record for anyone's time. Build cost for our own calendar would be a team of 8–10 for years — recurrence, tz, interop — to re-solve problems that aren't our product. See Buy or Build: The Total-Cost Test."

Why this is L6:

  • Locates the system of record correctly (the user's provider)
  • Names what to build (scheduling logic) vs integrate (calendar)
  • Plans for sync-token expiry as a normal event

What L7 adds:

  • Prices provider API quotas and the risk of one provider changing terms
  • Builds an internal availability abstraction so providers can be added or dropped without touching scheduling logic

Drill 8: Changing Policy Without an Outage — Shift Instead of Drop#

Prompt: "Today recurring instances in a DST gap are dropped, per the RFC. Product wants them shifted forward. How do you roll that out?"

Staff Answer

"This changes the meaning of existing series, so it's a data-semantics migration, not a code tweak. Step one: version the expansion policy — gap_policy=omit|shift — and stamp it on every series; existing series default to omit. Step two: shadow — compute both expansions for series with instances in gaps over the next 18 months and report how many differ; it's tiny (series at 02:00–02:59 local in DST zones), maybe tens of thousands. Step three: ship the new policy in every client behind a capability flag; the server only flips a series to shift once its attendees' clients support it, or renders server-side for those that don't. Step four: flip new series to shift, then migrate old ones, re-deriving busy intervals and reminders for each. Step five: tell organizers whose series gain an instance. iCalendar export stays RFC-correct: we emit an explicit RDATE for shifted instances so other systems see them too."

Why this is L6:

  • Treats semantic changes as migrations with a shadow phase
  • Coordinates server and client rollout through capabilities
  • Keeps interop correct by emitting explicit RDATEs

What L7 adds:

  • Writes the policy into the org-wide time standard so every product that expands recurrences follows it
  • Considers whether to propose the behavior upstream rather than diverge silently

Drill 9: Multi-Region and Data Residency#

Prompt: "EU customers require their calendar data to stay in the EU. Meetings span EU and US users."

Staff Answer

"Each calendar has a home region; truth for an event lives in the organizer's calendar's home region. An EU-organized meeting with US attendees: the event stays in the EU; US attendees' calendars hold a reference plus the minimum projection allowed by policy — times and busy status, and details only if the tenant's policy permits — replicated as an iTIP-style copy with SEQUENCE ordering. RSVPs from US attendees travel back to the EU as REPLY-style messages. Free/busy across regions reads each calendar's index in its home region; queries pay one cross-region hop (~80–120 ms transatlantic), which fits a 200 ms budget if fanned out in parallel. Reminders fire from the attendee's home region using the projected copy. During a region partition, each side keeps working on its own calendars; cross-region invites queue and apply in SEQUENCE order on heal. See Multi-Region."

Why this is L6:

  • Places truth by ownership and residency, not by user location
  • Reuses the iTIP ordering model internally across regions
  • States the latency cost and the partition behavior

What L7 adds:

  • Negotiates with legal what a "minimal projection" is, and makes it a tenant-visible setting
  • Prices cross-region replication and decides which regions get full stacks vs read-through

Drill 10: Cost#

Prompt: "Calendar infrastructure costs are up 40% year over year. Where do you look?"

Staff Answer

"Three suspects, in order. Sync: 'anything changed?' checks scale with connected clients times poll frequency, not with users; one client version polling every 60 s instead of 15 minutes multiplies that line by 15. I'd break down request rate by client version and app. Resyncs: full resyncs cost thousands of events each; if cal.sync.full_resyncs_per_min is up, find what's expiring tokens. Projections: busy-index write amplification from large events and series edits — one 60K-person event edited ten times is 600K index writes. Fixes: enforce client poll floors at the gateway, broadcast mode above 1,000 guests, lazy materialization for calendars that haven't been queried in 90 days (expand on read for them; most consumer calendars are never free/busy-queried by anyone else)."

Why this is L6:

  • Knows which load scales with clients rather than users
  • Ties cost to specific amplification paths with metrics
  • Proposes lazy materialization based on query patterns

What L7 adds:

  • Moves work to devices deliberately (local expansion, local alarms) and measures server savings
  • Sets per-partner API quotas so third-party clients carry their own cost

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Monday 08:50, Reminders Two Minutes Late#

Context: For the third Monday running, users in the largest region report 09:00 meeting reminders arriving at 08:52 or later. The timer service is "healthy": no errors, CPU at 40%. The VP of the enterprise product escalates to you after a customer's CEO missed a board call.

Questions to Surface First:

  • Is the lateness uniform, or concentrated in specific minutes of the hour?
  • Where does the time go — reading due timers, handing off to notifications, or delivery by the push provider?
  • Did anything change in timer partitioning or default reminder offsets recently?

Typical L5 Approach: Scales the timer worker fleet 3×. Lateness barely moves, because every worker is reading the same hot time-bucket partition.

Staff Approach: Breaks cal.reminder.lateness_p99 down by minute-of-hour and stage: lateness is only at :50 and :20, and it's all in the bucket read. The 08:50 bucket is one partition holding ~400K timers. Re-keys buckets as (minute, user_hash mod 64), adds pre-staging (load next minute 60 s early), and passes deliver_at to the notification system so delivery is on the second.

Principal Approach: Treats the top-of-hour herd as a capacity class across every time-triggered system in the company — reminders, scheduled reports, cron jobs — and sets a shared rule: time-indexed stores must sub-partition buckets and plan for the 99th-percentile minute. Also asks product whether a 9-minute default reminder offset for some users would flatten the peak without hurting experience.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm lateness by minute-of-hour; identify the hot bucket partition
TriageTimer bucket keyed by minute only; one partition per minute; read p99 4 s at :50
Quick fixPre-load the next three peak minutes into worker memory ahead of time for this week
GuardrailsSub-partitioned buckets; deliver_at hand-off; load test the :50 minute weekly
Post-mortemWhy was capacity planned on average rate? Why didn't the dashboard show lateness by minute-of-hour?

Metrics to Watch: cal.reminder.lateness_p99 by minute-of-hour, cal.timer.bucket_read_ms, notif.deliver_at_hold_ms

Organizational Follow-up: reminder lateness becomes an SLO with an error budget owned by the calendar platform; notification platform commits to burst capacity at :50.

Ownership Question: "Who owns a late reminder — calendar or notifications?" Staff answer: Calendar owns 'due timer handed off before deliver_at'; notifications own 'delivered by deliver_at + 10 s'. Each half has its own SLO, and the end-to-end probe pages calendar first.

Key Takeaway: "Calendar load is shaped by the clock on the wall. Plan for the minute everyone shares, not the day's average."

What clears the Staff bar:

  • Breaks lateness down by minute-of-hour and pipeline stage
  • Fixes partitioning, not fleet size
  • Splits the SLO between timer and delivery owners

Deep Dive 2: Silent Failure — Recurring Meetings an Hour Off in One Country#

Context: Support tickets trickle in from one country: some recurring meetings show an hour off, but only on some devices, and only for instances after last weekend. The country announced eleven days ago that it would not fall back this year. No alerts fired.

Questions to Surface First:

  • When did the tz release with this change come out, and when did we ingest it?
  • Which representation do affected series use — local + zone, or something else?
  • Which clients disagree with the server, and which tzdb version do they carry?

Typical L5 Approach: Updates the server's tz library and restarts services. The server now renders correctly; the busy index, room bookings and reminders computed before the update are still on old rules, and phones on old OS tz data still disagree.

Staff Approach: Finds that ingestion is manual and lagged the release by nine days. Runs the rebase for the changed zone (future instances only), re-derives reminders and room bookings in the same pass, and enables server-rendered times for clients in that zone whose reported tzdb version is older than the fix. Then automates ingestion with a 1-hour SLO and adds cal.tzdb.release_ingest_lag_hours as a paging alert.

Principal Approach: Makes tz data a tracked dependency for the whole company — every service and client reports the tzdb version it runs; a dashboard shows skew by platform; and a short-notice runbook (ingest, rebase, client fallback, customer comms) is owned by the calendar platform and exercised twice a year with a synthetic zone change.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm the zone and effective date; check server tzdb version
TriageIngestion manual; server 9 days behind; devices split by OS update status
Quick fixIngest release; rebase the zone; server-render for stale clients
GuardrailsAutomated ingestion; rebase lag alert; version-skew dashboard; reconciler checks a sample in each changed zone
Post-mortemWhy was a correctness dependency updated by hand? Why did no metric show server-vs-client disagreement?

Metrics to Watch: cal.tzdb.release_ingest_lag_hours, cal.tzdb.rebase_lag_hours, cal.tzdb.version_skew by zone and platform, cal.render.server_client_mismatch

Organizational Follow-up: tz ingestion owned by the calendar platform with an SLO; client teams own reporting their tzdb version.

Ownership Question: "Who decides what users see when server and phone disagree?" Staff answer: The calendar platform, by policy: for zones with a pending change, the server's computation wins and clients render it, because the server can be updated in an hour and phones can take weeks.

Key Takeaway: "Time zone data is a production dependency with a release cadence. If nobody owns ingesting it, it is a scheduled outage."

What clears the Staff bar:

  • Recognizes the fix spans index, reminders, rooms and devices — not just the server library
  • Rebases future instances only
  • Measures and mitigates version skew

Deep Dive 3: Large-Customer Onboarding — 400,000 Seats Migrating From Another Provider#

Context: A global enterprise is moving 400,000 users, 12,000 rooms and ~900 million events (including 15 years of history and ~40 million active recurring series) onto your platform over one weekend, with the old system still receiving edits until cut-over.

Questions to Surface First:

  • What's the import format — iCalendar, a proprietary API, or both? How are exceptions and time zones represented?
  • Which events are organized by people outside the migrating tenant?
  • How do we handle edits in the old system during migration?

Typical L5 Approach: Bulk-imports events through the public API over the weekend. Rate limits stretch it to nine days; recurring series with custom (non-IANA) time zone definitions import as fixed offsets; room double-bookings appear where the old system allowed them.

Staff Approach: Builds a bulk import path that writes intent directly to the event store per shard and lets projections catch up afterwards (with free/busy marked "warming" for the tenant). Maps every source time zone definition to an IANA ID — Windows zone names and custom VTIMEZONE blocks go through a mapping table, and unmappable ones are flagged for review rather than silently converted to fixed offsets. Imports history as frozen UTC, future events as intent. Runs a delta sync from the old system until cut-over. Imports rooms with the no-overlap constraint in report mode first and hands facilities a conflict list before enforcing.

Principal Approach: Turns this into a repeatable migration product: a published import contract, a zone-mapping service, and a tenant-level "migration mode" that relaxes quotas and defers projections. Prices it into enterprise deals, because every large customer will need it and each bespoke migration costs a team for a month.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Separate bulk path from public API; decide history vs future handling
TriageInventory zone definitions; count series with exceptions; find room conflicts
Quick fixZone mapping table with manual review queue; rooms imported in report mode
GuardrailsProjection warm-up with a visible "warming" state; delta sync until cut-over
Post-mortemWhich mappings were ambiguous? How long did projections take per million series?

Metrics to Watch: cal.import.events_per_s, cal.import.unmapped_tz_count, cal.projection.lag_s for the tenant, cal.room.overlap_count in report mode

Organizational Follow-up: migration tooling owned by a platform team, not the deal team; zone mapping maintained as shared data.

Ownership Question: "Who resolves the 3,000 room conflicts?" Staff answer: The customer's facilities team, with our conflict report, before enforcement is turned on. We don't silently pick winners for their rooms.

Key Takeaway: "A migration imports intent or it imports bugs. Fixed offsets are the bug."

What clears the Staff bar:

  • Distinguishes history (frozen UTC) from future (intent) at import
  • Refuses silent conversion of unknown zones to offsets
  • Enforces invariants in report mode first

Deep Dive 4: Post-Mortem — A "This and Following" Edit Deleted Three Months of Customizations#

Context: An executive assistant changed the time of a weekly leadership meeting "this and following". Afterwards, 14 instances that had been individually moved or had different rooms reverted to the series defaults; two board-prep sessions lost their rooms. The assistant files a complaint; you lead the post-mortem.

Questions to Surface First:

  • How does the split handle exceptions after the split point?
  • Did the UI warn that custom instances would change?
  • What happened to room bookings for re-homed instances?

Typical L5 Approach: Restores from backup for this one series and adds a confirmation dialog.

Staff Approach: Finds the split re-homed exceptions by original start, but the new series' rule generated different original starts (time changed from 10:00 to 11:00), so no exception matched and all were dropped. Fixes the split to re-key exceptions by mapping old original starts to new ones (same date, new time), preserving overridden fields; room bookings for re-homed instances are re-validated against the constraint and conflicts reported. Adds cal.series.exceptions_dropped and a pre-commit diff shown to the user: "14 customized instances — keep customizations / reset them".

Principal Approach: Treats series edits as a class of destructive operation that needs an undo: every series edit records a reversible change set for 30 days. Adds series-edit scenarios (with exceptions, rooms and external attendees) to the recurrence conformance suite every client runs.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Restore the series and exceptions from the change log, not backup
TriageException re-homing keyed on original start that no longer exists
Quick fixMap exceptions by date across the split; preserve overrides
GuardrailsPre-commit diff; cal.series.exceptions_dropped alert; 30-day undo
Post-mortemWhy could a single edit destroy data silently? Which other edit paths share this logic?

Metrics to Watch: cal.series.exceptions_dropped, cal.series.split_count, undo usage rate

Organizational Follow-up: recurrence edit semantics documented and owned by the calendar platform; clients may not implement their own split logic.

Ownership Question: "Who owns the definition of 'this and following'?" Staff answer: The calendar platform. It's a server-side operation with one implementation; clients call it, they don't reimplement it.

Key Takeaway: "Exceptions are user data. Any edit that can drop them must show a diff and support undo."

What clears the Staff bar:

  • Identifies the original-start keying as the root cause
  • Restores from the change log, not a backup
  • Adds a metric that makes silent data loss visible

Deep Dive 5: Multi-Region Expansion — Launching a Sovereign Region#

Context: The company is launching a region for public-sector customers who require that event data never leave the country. Their employees regularly meet with private-sector users in other regions.

Questions to Surface First:

  • What may leave the region — nothing, busy blocks only, or times without titles?
  • Who is the organizer of record for cross-region meetings?
  • How do reminders and push notifications work if the push provider is outside the country?

Typical L5 Approach: Deploys a full stack in the new region and disables cross-region invites.

Staff Approach: Full stack in-region with home-region calendars. Cross-region meetings use the iTIP-style copy model: when a sovereign user organizes, the event stays in-region and outside attendees receive only what policy allows (time, organizer, a generic title); when an outside user invites a sovereign user, the invite crosses as an iMIP-like message into the region. Free/busy across the boundary returns merged busy blocks only, rate-limited. Reminders fire in-region; push hand-off uses a provider path approved for the region, and email reminders use an in-country relay.

Principal Approach: Writes the data classification for calendar fields once — time, busy status, title, description, attendees, location — and defines which may cross which boundary, so every future region is a configuration of that table rather than a new design.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Agree the field-level crossing policy with legal
TriageIdentify all cross-region flows: invites, free/busy, reminders, search, notifications
Quick fixCross-region iTIP-style copies with field filtering
GuardrailsEgress audit on every cross-region message; synthetic tests that titles never leave
Post-mortem(Launch review) Which flows were discovered late?

Metrics to Watch: cal.xregion.messages_per_s, cal.xregion.policy_violations (must be 0), cross-region free/busy p99

Organizational Follow-up: data classification table owned by privacy engineering; calendar platform enforces it.

Ownership Question: "Who approves a new field crossing the boundary?" Staff answer: Privacy engineering and legal approve; the calendar platform implements and audits. Product cannot add a field to cross-region payloads on its own.

Key Takeaway: "Cross-region calendars are federation, not replication — the same iTIP model you already use with other companies."

What clears the Staff bar:

  • Reuses the iTIP copy model for cross-boundary meetings
  • Filters fields by policy, not by convenience
  • Covers reminders and notifications, not just storage

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why future events are stored as local time + IANA zone and past events as UTC, and what a tz release does to each
  • State DST gap and overlap behavior precisely, including the RFC 5545 defaults and the policy you'd choose
  • Model recurring series with exceptions keyed by original start, and implement this / this-and-following / all edits
  • Separate organizer-owned event data from per-attendee overlays, and use iTIP semantics at the boundary
  • Design a busy-interval index with a staleness SLO and privacy guarantees, and a no-overlap write path for rooms
  • Treat reminders as a rolling projection with SEQUENCE checks and plan for the :50 herd
  • Design change-log sync with tokens, push hints, poll backstops and paced full resyncs

The Bar for This Question#

Mid-level (L4): Builds events and attendees tables with UTC timestamps, a week view, and invites as copies. Recurring events are rows. Reminders are a cron query. Works in the demo; fails on the first DST change, series edit or offline device.

Senior (L5): Adds RRULE storage, attendees with RSVP status, a cache, sharding by user, and a reminder queue. Knows time zones exist and converts on display. The gap: UTC as the truth for future events, recurrence edits that lose exceptions, free/busy as a per-query expansion, rooms booked by check-then-write, and reminders created once. The design passes review and breaks on a government's announcement.

Staff+ (L6): Stores intent and treats instants as a cache, with a rebase pipeline tagged by tzdb version. Expands on read with exceptions keyed by original start. Splits organizer and attendee ownership. Answers free/busy from a projected index with a staleness SLO and protects rooms with a hard constraint. Makes reminders and sync re-derivable projections. Names who pays: users in changed zones during skew, organizers confirming destructive edits, the platform for rebase and reconciliation. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "Store Everything in UTC" Is Wrong for Calendars#

ClaimReality
"UTC is unambiguous"It is — and it's unambiguously wrong after a tz rule change for a future local-time commitment
"Convert on display"Display can't recover the zone the organizer meant once it's been discarded
"Rule changes are rare"tzdb shipped seven releases in 2022 and two-day notice is on record (tz NEWS)

The Staff position: UTC is the right representation for the past and for machine timestamps. For future human commitments, the truth is local time plus zone, and UTC is a derived cache.

Why this matters in interviews: "UTC everywhere" is the most common correct-sounding wrong answer in this question. Qualifying it is an immediate level signal.

10.2 The RFC's DST-Gap Default Is Not What Users Want#

The Staff position: RFC 5545 says recurring instances at nonexistent local times are ignored (RFC 5545). A user with a 02:30 daily series doesn't expect a missing day. Shift forward, document it, test it in every client, and emit explicit RDATEs at export so other systems agree.

Why this matters in interviews: Knowing the standard and deciding to deviate deliberately is stronger than either alone.

10.3 Most Calendars Should Not Have a Materialized Busy Index#

The Staff position: The index earns its keep for calendars someone else queries — enterprise users, rooms, booking hosts. Most consumer calendars are never free/busy-queried by anyone else. Materialize lazily on first query and expire after 90 days without one; expansion on read is microseconds per instance.

Why this matters in interviews: Knowing where not to precompute shows cost judgment, not just correctness.

10.4 Rooms Are Inventory, Not Attendees#

The Staff position: Treating rooms as attendees is elegant and wrong. A person decides whether to accept a conflict; a room must not have one. Rooms get the reservation-system invariant and a single write path.

Why this matters in interviews: It shows you can tell when two things that look alike in the UI have different invariants underneath.

10.5 Push Notifications for Sync Are a Hint, Never a Delivery Guarantee#

The Staff position: Google's own documentation says its calendar push notifications are not 100% reliable and carry no body (Google push guide). Any sync design whose correctness depends on pushes arriving is broken; the poll with a cursor is the guarantee.

Why this matters in interviews: It separates candidates who have integrated with real calendar APIs from those who drew a WebSocket and stopped.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a correct calendar. The Principal engineer notices that civil time is a company-wide dependency nobody owns. The calendar server has a tz library; the web client another; iOS and Android use the operating system's; the email digest renders times with a fourth; billing computes "end of month" in a fifth; the reminder and job schedulers each have their own DST policy. After every short-notice tz release these disagree for weeks, and the bug users report is never "wrong rule" — it is "my phone and my laptop show different times" or "my invoice is dated tomorrow". The L7 problem is not one calendar; it is one time model for the company: who ingests tz data and how fast, which recurrence engine everyone uses, which DST policy applies to people versus machines, and how skew is measured across every surface.

🧭 Principal Move: "Before I design the calendar, I want to know how many copies of the tz database we ship, how long it took each of them to pick up the last short-notice change, and how many recurrence expanders exist across our clients. The calendar's correctness is bounded by the slowest of those, so that's where I'd invest first."

The Org-Level Fault Line#

One time-and-recurrence platform vs each product owning its own.

OptionWhat WorksWhat BreaksWho Pays
Each product and client owns its time handlingAutonomy; fits each platformN tz versions, N RRULE edge-case behaviors, N DST policies; skew after every releaseUsers (inconsistency); support; every post-mortem
One central time service everyone callsOne truthNetwork call on every render; mobile offline breaksLatency; offline users
Platform owns the contract, data and conformance suite; products embed a certified engine (Staff/L7 default)Same behavior everywhere, works offline; tz updates shipped as dataConformance suite and data pipeline must be maintained; clients must updatePlatform (stewardship); client teams (adoption)

🧭 Principal Move: "The platform owns what must be identical everywhere — tz data ingestion, the recurrence semantics, the DST policy for people versus machines, and the conformance suite. Products own rendering and UX. Any client that can't pass the suite renders server-computed times. That's how we make a two-day-notice change a data push instead of five incidents."

Cost Model#

Assumptions: fully loaded engineer ~$250K/year, cloud list prices, rough ranges (estimates, not quotes).

ScaleUsers / SeatsInfra ($/month)HeadcountOn-call LoadBuy Alternative
Startup feature (scheduling inside another product)100K users~$1–3K (integrations, cache)1–2 eng on provider integrationsLow; provider outages are the pagesIntegrate with existing providers — always
Growth (own calendar for a vertical, e.g. clinics)5M users, 200K hosts~$30–80K (event store, index, timers, sync)8–12 eng: core, sync/interop, mobile, bookingShared rotation; tz releases and Monday peaks are the pagesPartial: buy booking UI, own data
Planet-scale500M users, 60M seats~$3–8M (sync and reminder fan-out dominate)150–300 eng across core, interop, clients, rooms, enterprise adminDedicated rotations per tier; correctness SLOsNot applicable

The pricing insight: storage is never the cost driver. Connected clients times poll frequency, reminder peak capacity, and projection fan-out from large events are. Every reminder a phone fires locally and every week view it expands locally is server capacity you don't buy — so the client architecture is a cost decision.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Storing future events as local + zone vs UTCOne-wayOnce intent is discarded it can't be recovered; migrating back means guessing zones for billions of events
Instance identity = (event, original start)One-wayEvery exception, RSVP, reminder and external client keys off it
Recurrence semantics (DST gap policy, invalid dates)One-way-ishChanging them moves existing instances for millions of series; needs a versioned migration (Drill 8)
iCalendar UID namespace for exported eventsOne-wayExternal systems keep UIDs forever; reuse or change creates duplicates
Organizer-owned event vs per-attendee copiesOne-way-ishRe-modelling invites touches every write path and every interop adapter
Busy-index horizon (18 months)Two-wayA projection; rebuild with a new horizon
Event-store technologyTwo-wayBehind the event service; dual-write and migrate per shard
Push vs poll mix for syncTwo-wayClient and gateway config

🧭 Principal Insight: The time representation and instance identity are the decisions I'd slow down on. Stores, horizons and sync cadences can change in a quarter; "what does this event mean" and "which instance is this" cannot change at all once billions of events and every external client depend on them.

The Standard I'd Write#

RFC-TIME-001: Civil Time and Recurrence Standard
Status: Approved   Owners: Calendar Platform + Client Platform Leads

Scope
  Every service and client that stores, schedules, renders or bills against
  human (civil) time: calendar, reminders, booking, notifications quiet hours,
  billing periods, scheduled reports.

MUST
  1. Store future human commitments as local time + IANA zone ID (or date, or
     floating time), never as a UTC instant or fixed offset alone.
  2. Store past events and machine timestamps as UTC instants.
  3. Tag every derived UTC instant with the tzdb version used to compute it.
  4. Ingest each IANA tz release within 1 hour of publication and complete
     rebase of derived future instants within 6 hours.
  5. Expand recurrences only with the certified engine, or a client that passes
     the conformance suite for the current semantics version.
  6. Apply the published DST policy: people-facing events shift out of gaps
     and take the first occurrence in overlaps; machine jobs follow the job
     scheduler's policy, documented separately.
  7. Report the tzdb version in use from every client and service.

SHOULD
  1. Render server-computed times when the client's tzdb is older than the
     server's for the zone being displayed.
  2. Emit explicit RDATE/EXDATE at iCalendar export wherever internal policy
     differs from RFC 5545 defaults.

Exceptions
  Filed with Calendar Platform; reviewed within 5 business days; time-boxed to
  two quarters.

Success metrics
  - cal.tzdb.release_ingest_lag_hours p99 ≤ 1
  - cal.tzdb.rebase_lag_hours p99 ≤ 6
  - cal.render.server_client_mismatch ≤ 0.01% of rendered instances
  - Recurrence engines in production: 1 (plus certified ports)

What I'd Tell the VP#

"Calendars break on a schedule we don't control: governments change clock rules, sometimes with two days' notice, several times a year. Today, five parts of our product handle time separately, and each takes a different amount of time to catch up — so after every change, customers see different meeting times on their phone and their laptop. I'm proposing one owner for time handling, an automated pipeline that applies rule changes within hours, and a single recurrence engine every app uses. It's roughly six engineers for a year, mostly consolidating work we already do in five places. Success is measured by how fast we apply changes and how often our apps disagree, which I'll report quarterly. The cost is that client teams give up their own time code, and I'll make sure they get a library that's better than what they have."

Principal Interview Signals#

SignalWhat It Sounds Like
Counts time implementations, not events"How many copies of tz data and how many RRULE expanders do we ship?"
Prices the client architecture"Every reminder the phone fires locally is :50 capacity we don't buy."
Sets the org's correctness posture"Rebase within six hours of a tz release is an SLO with an owner and a game day."
Identifies the real one-way doors"Local-plus-zone and instance identity are forever; the store isn't."
Knows when not to centralize"Rendering and UX stay with products; semantics and data are the platform's."

Staff answers that L7 interviewers find insufficient:

  • "We store local time plus zone and rebase on tz changes" — correct for the calendar; silent on the other four surfaces that render or schedule the same events.
  • "Clients use a standard RRULE library" — which one, at what version, tested against what?
  • "We'll support CalDAV and Exchange" — no statement of which is first-class, who owns it, or what it costs per enterprise deal.

Appendices

Appendix A: Mechanics in Depth#

A.1 Range Expansion#

def instances(series, window_utc, tzdb):
    z = tzdb.zone(series.tzid)                         # IANA rules for this zone
    lo = z.to_local(window_utc.start) - series.duration  # widen for overlap
    hi = z.to_local(window_utc.end)
    out = []
    for local in rrule_iter(series.rrule, series.start_local, lo, hi):
        if not valid_date(local):                      # Feb 30, Apr 31
            continue                                   # RFC default: omit
        utc = z.resolve(local, gap="shift_forward", overlap="first")
        orig = (series.id, local)                      # identity: original start
        if local in series.exdates: continue
        exc = exceptions.get(orig)
        inst = apply(exc, series, utc) if exc else make(series, utc)
        if overlaps(inst, window_utc): out.append(inst)
    out += [make(series, z.resolve(d)) for d in series.rdates if in_window(d)]
    out += exceptions.moved_into(series.id, window_utc)  # by current start
    return dedupe_by_original_start(out)
def free_slots(calendars, window, duration, working_hours):
    busy = []
    for cal in calendars:                               # parallel, one shard each
        busy += busy_index.scan(cal, window)            # [start_utc, end_utc)
        busy += outside_working_hours(cal, window)      # in each attendee's zone
    busy.sort()
    merged = merge_overlapping(busy)                    # O(n log n)
    return [gap for gap in gaps(merged, window) if gap.length >= duration]

Working hours are wall-clock ranges in each attendee's zone, so they are converted per day — a 09:00–17:00 Berlin day and a 09:00–17:00 New York day overlap by a different number of hours in the March weeks when only one side has switched.

A.3 TZ Rebase#

Diagram: A.3 TZ Rebase

Room bookings are re-derived inside the no-overlap constraint; a rebase can create a conflict when a booking anchored in a changed zone moves onto one anchored in an unchanged zone (a London-organized meeting in a room also booked from a Cairo-anchored series, say). Conflicts are reported to the organizers, never silently resolved.

Appendix B: Data Model#

calendars(calendar_id PK, owner_kind, owner_id, default_tzid, home_region)
events(event_id PK, calendar_id, ical_uid UNIQUE, start_local, duration,
       tzid NULL,                 -- NULL = floating
       all_day bool, start_date, end_date,          -- dates for all-day, half-open
       rrule NULL, rdates[], exdates[], until_utc,
       sequence int, split_from NULL, title, location, visibility, status)
exceptions(event_id, original_start_local, overrides_json, cancelled,
           current_start_utc, PK(event_id, original_start_local))
attendance(event_id, attendee_kind, attendee_id, partstat, role,
           personal_reminders_json, hidden, last_seen_sequence,
           PK(event_id, attendee_kind, attendee_id))
busy_intervals(calendar_id, start_utc, end_utc, event_id, original_start_local,
               transparency, tzdb_version)          -- projection, 18 months
room_bookings(room_id, during tstzrange, event_id, original_start_local)
               -- EXCLUDE USING gist (room_id WITH =, during WITH &&)
reminder_timers(bucket_minute, user_hash_mod_64, user_id, event_id,
                original_start_local, offset_min, channel, sequence)
change_log(calendar_id, seq, event_id, op, at)      -- 30-day retention

The shard key is calendar_id for everything owned by a calendar; attendance rows live with the event (organizer's shard) and are projected into each attendee's change log and busy index on their own shards. See Schema Design for the access-pattern-first approach and Partitioning for co-location.

Appendix C: Coordination Mechanisms#

MechanismUsed ForWhy Not Something Else
Single-shard transaction on organizer's calendarEvent + overlays + change log appendWrites stay local; attendees' views are projections
Optimistic concurrency on SEQUENCE (If-Match)Concurrent organizer edits (assistant + exec)Conflicts are rare and must be shown to humans
Exclusion constraint on room rangesRoom and booking-page slotsThe store enforces the invariant; no lock service
Idempotency keys on bookings and RSVPsRetries from mobile and email gatewaysSee Idempotency
Change stream (Kafka, keyed by event_id)Projections in order per eventPer-event ordering is enough; no global order needed
iTIP SEQUENCE orderingExternal and cross-region copiesThe standard already defines it

Appendix D: Sync Protocol and Client Behavior#

Diagram: Appendix D: Sync Protocol and Client Behavior
  • Tokens are per calendar, encode (calendar_id, seq), and are valid for 30 days of change-log retention; older tokens get 410 with a paced Retry-After.
  • Push is a hint with no payload; the device always pulls with its token. Channels expire and are renewed by the client.
  • External CalDAV clients use the RFC 6578 sync-collection REPORT against the same change log, and receive whole iCalendar objects per changed resource; the adapter must round-trip RRULE, exceptions and VTIMEZONE without loss.
  • Offline edits carry the SEQUENCE the device last saw; the server rejects stale organizer edits with a conflict the user resolves, and always accepts RSVP overlay writes (last writer wins on the attendee's own row).
  • Local alarms: devices schedule 7 days of alarms from synced data and report the SEQUENCE scheduled per event so the server can suppress duplicate pushes.

Appendix E: Observability#

MetricAlertWhy
cal.tzdb.release_ingest_lag_hours> 1 hCorrectness dependency is behind
cal.tzdb.rebase_lag_hours> 6 hFuture instances still on old rules
cal.tzdb.version_skew by zone and platformWeekly report; page on changed zonesDevices disagree with the server
cal.reminder.lateness_p99 by minute-of-hour> 30 sThe :50 herd or a projection lag
cal.freebusy.staleness_s p99> 5 sPeople double-book what they saw free
cal.freebusy.divergence (reconciler)> 0.1%Silent projection bugs
cal.room.overlap_count> 0A side-door write bypassed the constraint
cal.sync.full_resyncs_per_min> 3× baselineToken invalidation storm
cal.series.exceptions_droppedAny spikeDestructive series edits
cal.render.server_client_mismatch> 0.01%Recurrence or tz disagreement across clients

Correctness vs availability: most calendar alerts here are correctness signals that would never show up as error rates. Page on rebase lag, divergence and room overlaps at all hours; they are the incidents users experience as "I was in the wrong place at the wrong time".

Appendix F: Scale Evolution#

StageWhat WorksWhat You Add
< 100K usersIntegrate with existing providers; no calendar of your ownShort-TTL availability cache; sync-token handling
100K–5MOwn calendar on one PostgreSQL cluster; expand on read; timers in a queueIntent-based time model from day one; room constraint
5M–100MSharded by calendar; busy index; change-log sync; projectionstzdb automation; broadcast events; paced resync
100M+Multi-region home placement; device-side expansion and alarmsCompany-wide time standard; conformance suite; skew dashboards

What you don't build on day one: your own tz database, a custom recurrence format, smart time suggestions, cross-tenant federation, or a materialized busy index for calendars nobody queries.

  1. Loading the index…