Technologies referenced in this case study: PostgreSQL · Cassandra · Redis · Kafka · Distributed SQL
Related: Job Scheduler · Notifications · Reservation Systems · Cloud File Sync · Schema Design · Hot Keys · Real-Time Updates · Idempotency
Reading Guide#
Organized for interview use first, reference second. This page designs a calendar service: events, recurring series, invitations and RSVPs, free/busy across many people, rooms, booking pages, reminders, and sync with clients the service does not control. Three neighbours own pieces of the machinery and this page links to them instead of repeating them: firing a timer reliably at a wall-clock moment is Job Scheduler; getting a push or email to a phone is Notifications; never selling one room twice is Reservation Systems. Cursor-based sync with offline clients is covered in depth in Cloud File Sync.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and the appendices on recurrence expansion, the time model and sync |
What is a Calendar & Scheduling Service? — Why interviewers pick this topic
A calendar service stores commitments in civil time — "every Tuesday at 09:30 in London", "the board meeting on the last Thursday of each quarter", "Priya's birthday on 29 February" — and answers three questions about them, millions of times a second: what is on my calendar this week, when are these eight people and one room all free, and what should ring on my phone in ten minutes. On top of that it moves invitations between people (inside the company and outside it), keeps phones, laptops and third-party clients in sync, and lets strangers book a slot on a booking page.
The hard part is not storing a row with a start and an end. The hard part is that most of the events users care about have not happened yet, and the meaning of a future time can change after you store it. A weekly meeting has no last instance. "09:30 London" maps to a different UTC instant in July than in January. Governments change daylight-saving rules with days of notice, and every future meeting in that zone must move — or must not move, depending on which zone the organizer meant. One edit to a series with 300 attendees touches 300 calendars, thousands of reminder timers, and every device those people own.
Before vs After — the "government changes the clocks" scenario:
Without a designed time model:
t=-2y: Events stored as UTC instants. A weekly 09:00 Mexico City stand-up is
expanded into 104 rows of UTC timestamps, assuming DST continues.
t=0: Mexico abolishes DST for most of the country. New tz data ships.
t=+1day: Nothing recomputes. Stored instants are now one hour off for every
winter instance. Phones (with new OS tz data) render 08:00 for some
users and 09:00 for others, depending on whether the client re-derives.
t=+1week: First Monday after the change: half the team joins an hour early.
Room bookings and free/busy now disagree with what users see.
t=+3weeks: Support finds 4M affected series. Fix requires a migration over
data whose original intent ("09:00 Mexico City") was never stored.
With a designed time model:
t=0: Events stored as (local start 09:00, zone America/Mexico_City, RRULE).
Derived UTC instants in the index carry the tzdb version used.
t=+1h: New tzdb release ingested; diff shows which zones changed from which date.
t=+2h: Rebase job recomputes future index entries and reminders for affected
zones only. Past instances are untouched.
t=+3h: Every client and the free/busy index agree: 09:00 local, new offset.
Cross-zone attendees see the meeting move in their own zone — correctly.
Why interviewers reach for this question: It looks like CRUD with a nice UI, so it cleanly separates candidates who model rows from candidates who model intent. The traps are all correctness traps that pass every demo: storing UTC for future events, materializing infinite series, treating the organizer's and the attendee's copies as the same object, computing free/busy by scanning events, and scheduling reminders once at creation time. Each one fails months later, on a DST weekend or a policy change, for millions of users at once.
Mechanics Refresher: The Calendar Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| UTC instant | 2026-11-03T14:00:00Z | Unambiguous; sorts and compares trivially | Loses intent: "09:00 in Chicago" cannot be recovered after a rule change |
| Wall-clock + IANA zone | 09:00 + America/Chicago | Preserves what the user meant; survives tz rule changes | Must be converted to compare; DST gaps and overlaps need a policy |
| Floating time | 09:00 with no zone | "Take medication at 08:00 wherever I am" | Means a different instant for every viewer; cannot go in a shared free/busy index |
| All-day / date value | 2026-12-25 (a date, not an instant) | Holidays and birthdays don't drift across zones | Spans different UTC ranges per viewer; must not be stored as midnight UTC |
| RRULE | FREQ=WEEKLY;BYDAY=TU;INTERVAL=1 from RFC 5545 | One row describes an infinite series | Must be expanded to answer any time-range question |
| Exception (RECURRENCE-ID) | An override keyed by the instance's original start | Moves or edits one instance without touching the rule | Orphaned when the rule changes underneath it |
| EXDATE | List of cancelled original starts | Cancels single instances cheaply | Same orphaning risk; lists grow over years |
| Series split ("this and following") | End the old rule with UNTIL, start a new series | Clean history; past instances keep old details | Two series to keep linked; exceptions must be re-homed |
| Free/busy | Busy intervals only, no titles | Privacy-preserving availability across people | Must reflect expanded recurrences and every RSVP |
| iTIP message | REQUEST / REPLY / CANCEL with UID + SEQUENCE (RFC 5546) | Interoperable invites across systems | Ordering and duplicates are the receiver's problem |
| Sync token | Opaque cursor into a calendar's change log | Clients fetch only deltas | Expiry forces a full resync; a storm if many expire at once |
For most production systems: store future events as wall-clock + IANA zone + RRULE + exceptions (the intent), derive UTC instants into a bounded-horizon index stamped with the tzdb version (the cache), keep one shared event with per-attendee overlay rows inside the system and iTIP at the boundary, answer free/busy from the busy-interval index, and materialize reminder timers on a rolling horizon so edits and tz changes re-derive them. The primitives are not the interview — what each stored value means when the rules change is.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
A calendar is not a CRUD question. Anyone can store an event with a start and an end.
It is a time-and-recurrence correctness question that tests:
- Whether you store what the user meant (a wall-clock time in a named zone, and a rule) rather than what it happened to mean on the day you stored it
- Whether you can answer "what is on these calendars between T1 and T2" when most events are infinite rules plus exceptions
- Whether you see that one event has many owners — the organizer owns the time, each attendee owns their RSVP, reminders and colour — and design the write path around that split
- Whether you treat reminders, free/busy and device sync as derived views that must be re-derived when the rule, the RSVP or the tz database changes
The key insight: For a calendar, the source of truth for a future event is intent, and every instant is a cache. Past events are facts in UTC; future events are wall-clock times in a zone, governed by rules that both users and governments can change. Staff candidates say this in the first five minutes and then design the free/busy index, the reminder timers and the sync change log as re-derivable from intent — with an explicit job that re-derives them when the tz database changes.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws events table → API → clients; adds recurring events as a flag | Asks "Personal, enterprise scheduling or booking pages? Which zone anchors a meeting? How far ahead must free/busy be exact, and which clients do we not control?" | Asks "How many time libraries and tz data copies do our server, web, iOS, Android and email stacks embed, and how do we know they agree?" |
| Time | "Store everything in UTC, convert on display" | "UTC for the past, wall-clock + IANA zone for the future. Derived instants carry a tzdb version; a tz release triggers a rebase of affected future instances." | Owns tz data as a company-wide dependency: one ingestion pipeline, a skew SLO across clients, and a short-notice-change runbook |
| Recurrence | "Generate all instances into the events table" | "Store the RRULE and exceptions keyed by original start; expand on read; materialize only a bounded horizon into the busy index" | Publishes one recurrence engine with a conformance suite every client must pass; RRULE semantics become a one-way-door contract |
| Invites | "Copy the event into each attendee's calendar" | "One shared event owned by the organizer, a per-attendee overlay for RSVP and reminders, iTIP with SEQUENCE at the boundary for external attendees" | Decides interop posture: which external protocols (CalDAV, iMIP, Exchange) are first-class, and who funds their long tail |
| Free/busy | "Query each attendee's events for the range" | "A per-calendar busy-interval index, 18 months materialized, merged per query; privacy enforced by returning intervals only" | Treats availability as a product API with a published freshness SLO that rooms, booking pages and partner tools build on |
| Scale | "Shard events by user" | "Shard by calendar; the hot reads are sync 'anything changed?' checks (~450K/s) and the hot spike is reminders at :50 past the hour" | Prices it: sync and reminder fan-out dominate cost; decides what runs on devices (local alarms, expansion) to cut server load |
Why "time" separates levels
L5: "Everything is stored in UTC; the client converts to local time for display." This is correct for logs, payments and past events, and it is the standard advice for most backends. For a future recurring meeting it silently destroys information: once "09:00 in Chicago every Tuesday" is turned into a list of UTC instants, a change to Chicago's DST rules cannot be applied, and you cannot even tell which events were meant to follow Chicago rather than the UTC offset they happened to have.
L6: "A future event is stored as a local start, a duration and an IANA zone ID — America/Chicago, never CST or -06:00. I derive UTC instants into the index for range queries, and every derived row records the tzdb version it was computed with. When a new tzdb release lands, I diff it, find zones whose future offsets changed, and recompute only those future instances, their busy intervals and their reminders. Past instances freeze as UTC: what happened, happened. All-day events are dates, and 'take my pill at 08:00' is floating time with no zone at all."
L7: "Our server uses one tz library, the web client another, iOS and Android use whatever the OS ships, and the email renderer has its own. After a short-notice change those disagree for weeks. I'd make tz data a platform dependency with an owner, an ingestion SLO measured in hours, and a dashboard of tzdb versions observed across clients — because the bug users see is not 'wrong time', it's 'my phone and my laptop disagree'."
Why "recurrence" separates levels
L5: "When someone creates a weekly meeting, we insert a row per instance for the next two years." Reasonable for a demo. Then the organizer changes the time: 104 rows rewritten, exceptions lost, and the series silently ends two years from creation. A daily standup with 40 attendees materialized for two years is ~29,000 attendee-instance rows from one click.
L6: "The series is one row with an RRULE. Exceptions are separate rows keyed by (series_id, original_start) — the RFC 5545 RECURRENCE-ID — so they survive edits to other instances. Range queries expand the rule on read; it's microseconds per instance. The only thing I materialize is busy intervals for the next 18 months, because free/busy across 20 people needs an index, not 20 rule expansions — and that materialization is a cache I can rebuild."
L7: "There are five RRULE expanders in this company — server, web, two mobile apps and the email digest — and each disagrees with the others on at least one edge case: BYSETPOS, the 31st of short months, DST gaps. I'd ship one engine, compiled for every platform or behind one service, with a shared conformance suite, and I'd treat its semantics as a contract we never change silently."
Why "invites" separate levels
L5: "Each attendee gets their own copy of the event in their calendar." It makes reads simple and it makes every edit a fan-out write to N copies that can disagree. A 300-person all-hands moved by 30 minutes is 300 writes, 300 sync notifications and, if one copy fails, one person who shows up at the old time.
L6: "Inside our system there is one event, owned by the organizer's calendar. Each attendee has an overlay row: RSVP status, personal reminders, colour, whether it's hidden. Time, title and rule live only on the event, so an edit is one write plus index and notification fan-out. External attendees get iTIP messages over email — REQUEST, REPLY, CANCEL — and their SEQUENCE numbers tell both sides which version wins."
L7: "Interop is where this product's cost hides: Exchange, CalDAV servers, ICS feeds, mail clients that render invites their own way. I'd decide explicitly which protocols get first-class support and an owning team, and which are best-effort — because otherwise every enterprise deal adds one more half-supported integration."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Future events are wall-clock + IANA zone; past events are UTC instants | Governments change rules with days of notice; only the intent survives that. The past cannot change, so freeze it |
| Store the rule, expand on read, materialize a bounded horizon | Infinite series cannot be rows; free/busy needs an index; 18 months covers almost all scheduling while staying rebuildable |
| Exceptions are keyed by original start, never by current start | RECURRENCE-ID is what keeps a moved instance attached to its series across edits |
| One shared event, per-attendee overlays; iTIP only at the boundary | An edit is one write, not N; RSVPs never contend with the organizer's edits |
| Free/busy is a derived busy-interval index, not a query over events | 8 attendees × rule expansion per query does not survive 80K queries/s; the index does |
| Reminders are a rolling-horizon projection, re-derived on every change | Scheduling at creation time goes stale on the first edit, RSVP decline or tz release |
| Sync is a per-calendar change log with tokens; push is a hint, poll is the guarantee | Clients you don't control miss pushes; tokens that expire must degrade to a resync, not data loss |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Personal & shared calendars (consumer) | 500M users, many devices each, offline edits, third-party clients | Intent-based time model, expand-on-read, change-log sync, device-local alarms plus server reminders | Wrong time after a DST or tz rule change; devices disagree; missed or duplicate reminders | Every client renders the same local time for an instance; reminder on time p99 ≤ 30 s |
| Enterprise scheduling (free/busy, rooms, delegation) | 60M seats, meetings across zones and companies, rooms as contended resources, privacy rules | Busy-interval index, organizer-owned events with attendee overlays, room booking as a conditional write, iTIP with external systems | Double-booked rooms; free/busy stale after edits; private details leaking through availability | Free/busy reflects writes within 5 s; zero double-booked rooms; no title leaks to non-delegates |
| Booking pages / appointment scheduling | Strangers book slots in a host's availability; bursts when a link is shared; payments sometimes attached | Availability = working hours − busy − buffers, computed from the same index; booking as an idempotent conditional insert | Two bookers get the same slot; host's new meeting races a booking; slot shown in the wrong zone | Zero double bookings; slot times shown in the booker's zone and stored in the host's |
🎯 Staff Move: "I'll design the shared core first — the time model, recurrence, the organizer-owned event and the busy-interval index — because personal calendars and enterprise scheduling both stand on it. I'll go deepest on enterprise scheduling: free/busy across many people, rooms and invites are where the correctness bar is highest. Booking pages reuse the same availability computation with a reservation-style write path, which I'll cover at the end."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Recurrence: Store the Rule vs Materialize Instances | Compact, editable, infinite series vs simple range queries and indexable instances |
| 2 | Time Representation: UTC Instant vs Wall-Clock + Zone | Trivial comparison and sorting vs surviving tz rule changes and DST without losing intent |
| 3 | Invitations: Shared Event vs Per-Attendee Copies | One write per edit and one truth vs independent copies that read cheaply and work across systems |
| 4 | Free/Busy: Compute on Read vs Precomputed Busy Index | Always-correct, no extra state vs fast multi-person queries with an index that can go stale |
| 5 | Reminders: Schedule at Write vs Rolling-Horizon Projection | Simple one-time timers vs reminders that follow edits, RSVPs and tz changes |
How Real Companies Built It#
Why this section belongs here: Calendars run on public standards and public data. Every fault line on this page is visible in an RFC, the tz database's release history, or a large provider's documented API behavior.
RFC 5545 — iCalendar Recurrence Has Sharp Edges by Design#
RFC 5545 (September 2009) defines the RRULE grammar every major calendar speaks, and it settles three edge cases on paper: an RRULE that generates an invalid date (30 February) or a nonexistent local time (inside a DST spring-forward gap) MUST ignore that instance and not count it toward COUNT; a DATE-TIME whose local time occurs twice in a fall-back overlap refers to the first occurrence; and a DTSTART inside the gap is interpreted with the offset from before the gap, so America/New_York 02:30 on 11 March 2007 means 03:30 EDT. It also defines RECURRENCE-ID with RANGE=THISANDFUTURE, and forbids COUNT and UNTIL in the same rule (RFC 5545). RFC 7529 (May 2015) later added RSCALE for non-Gregorian calendars and a SKIP option — OMIT (the default), BACKWARD or FORWARD — so a 29 February birthday can land on 1 March in non-leap years instead of vanishing (RFC 7529).
Staff insight: The standard's default for a "monthly on the 31st" rule is to skip the seven months without a 31st, and its default for a recurring 02:30 meeting on spring-forward day is to drop that instance, while a single event at 02:30 is moved to 03:30. Users rarely expect either. Name the policy you implement, implement it identically on every client, and say so in the interview — it is the fastest way to show you have actually expanded RRULEs.
IANA Time Zone Database — Rules Change on Days of Notice#
The tz database records "the history of local time for many representative locations" and is updated when political bodies change offsets or DST rules; RFC 6557 (BCP 175, February 2012) makes the TZ Coordinator an IANA Designated Expert who decides on changes and release timing (IANA, RFC 6557). The release history shows the cadence: seven releases in 2022 (2022a–g), four in 2023, two in 2024, three in 2025, and five by the end of September 2026. Release 2022f, published 28 October 2022, recorded that Mexico would stop observing DST after 2022 (except near the US border) and that Chihuahua would move to year-round −06 on 30 October — two days later. Release 2023b (23 March 2023) moved Lebanon's spring-forward from 25/26 March to 20/21 April; 2023c, five days later, reverted it. In 2026, release 2026d recorded that Canada's Northwest Territories would not fall back on 1 November 2026, and 2026e (29 September 2026) that Manitoba stays on −05 (tz NEWS).
Staff insight: A calendar is the system most exposed to these releases, because it holds future local times. Two-day notice means the rebase of future instances, busy intervals and reminders must be an automated pipeline measured in hours, and that server and devices will disagree until OS vendors ship the same data. Design for skew, not for agreement.
CalDAV, iTIP and iMIP — Sync and Scheduling as Standards#
CalDAV (RFC 4791, March 2007) extends WebDAV with a calendar-query REPORT with time-range filtering and optional expansion of recurring events, calendar-multiget, and a free-busy-query REPORT that returns availability without event details (RFC 4791). iTIP (RFC 5546, December 2009) defines the scheduling methods — PUBLISH, REQUEST, REPLY, ADD, CANCEL, REFRESH, COUNTER, DECLINECOUNTER — with UID as the primary key and an organizer-incremented SEQUENCE so a higher sequence obsoletes earlier versions (RFC 5546); iMIP (RFC 6047) carries those messages over email (RFC 6047). CalDAV Scheduling (RFC 6638, June 2012) moves iTIP delivery to the server ("implicit scheduling") with scheduling inbox and outbox collections (RFC 6638), and WebDAV sync (RFC 6578, March 2012) adds sync tokens, with a DAV:valid-sync-token error that sends clients back to a full sync when the server no longer has the history (RFC 6578).
Staff insight: The standards already encode the Staff answers: the organizer owns the event and SEQUENCE orders versions; free/busy is a separate, detail-free query; sync is a cursor that can expire. If you support CalDAV you inherit clients that poll on their own schedule and send whole iCalendar objects back — so your internal model must round-trip iCalendar without losing exceptions.
Google Calendar API — Sync Tokens, Bodiless Pushes, Bounded Free/Busy#
Google's Calendar API documents incremental sync with nextSyncToken, present only on the last page of a full sync; when a token is invalidated — "token expiration or changes in related ACLs" — the server returns 410 and the client should wipe its store and do a full sync (Google sync guide). Push notifications carry no body, only headers such as X-Goog-Resource-State; channels expire and must be replaced by calling watch again; and the guide warns that notifications "are not 100% reliable" and a small percentage are dropped (Google push guide). The free/busy query accepts calendarExpansionMax up to 50 and groupExpansionMax up to 100 and returns only busy ranges per calendar (freebusy.query). Recurring-event instances carry recurringEventId and originalStartTime (Google recurring events guide).
Staff insight: This is the Staff design as an API contract: push is a hint and polling with a cursor is the guarantee; an expired cursor costs a full resync, so cursor retention is a capacity decision; free/busy is bounded per request so one query cannot fan out to an entire company; and instances are identified by their original start. Cite the 410 behavior when asked what happens to offline clients.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Store times in UTC" | "Chile changes its DST date next week. What happens to a recurring 09:00 meeting there?" | Intent vs instant; tz rebase |
| "Recurring events are expanded into rows" | "The organizer moves the series 30 minutes later. What happens to the instance someone already moved?" | Exceptions keyed by original start; series edits |
| "Each attendee gets a copy" | "The organizer edits the title while an attendee declines. Which write wins?" | Ownership split; overlay rows; SEQUENCE |
| "We query attendees' calendars for free/busy" | "Find a 30-minute slot for 12 people and a room next week. How many reads is that?" | Busy-interval index; merge cost |
| "We schedule a reminder when the event is created" | "The attendee declines, then the organizer moves it. Does the reminder still fire?" | Reminders as a re-derived projection |
| "Clients sync via the API" | "An iPhone has been offline for six weeks. What does it fetch when it reconnects?" | Change-log retention; token expiry; full resync cost |
| "Rooms are attendees" | "Two people book Room 4B for 10:00 at the same moment. Who gets it?" | Conditional write on the room's busy set |
System Architecture Overview#
Reading the diagram: The event store holds intent — local times, zones, rules, exceptions — and is the only source of truth. Everything to its right is derived: the busy-interval index (UTC instants for the next 18 months), the reminder timers (next 48 hours), the per-calendar change logs that clients sync from, and outbound iTIP messages. Derivation runs off a change stream, so an edit is one transactional write plus asynchronous projection. The TZ rebase job is the piece most designs omit: it re-derives future instants when the tz database changes. Reminder delivery hands off to Notifications; timer firing follows Job Scheduler.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Time | "Store UTC" | "Past in UTC. Future as local time + IANA zone; derived instants carry a tzdb version; tz releases trigger a rebase." |
| Recurrence | "Insert a row per instance" | "Store RRULE + exceptions keyed by original start; expand on read; materialize busy intervals 18 months out." |
| Invites | "Copy to each attendee" | "One organizer-owned event, per-attendee overlay rows; iTIP with SEQUENCE for external attendees." |
| Free/busy | "Query each calendar" | "Busy-interval index per calendar, merged per query; intervals only, no titles; ≤ 5 s staleness SLO." |
| Rooms | "Rooms are attendees" | "Rooms are resources with a no-overlap constraint; booking is a conditional insert, not a check-then-write." |
| Reminders | "Schedule at create" | "Rolling 48 h timer projection re-derived on edit, RSVP and tz change; dedupe key includes SEQUENCE." |
| Sync | "Clients call the API" | "Per-calendar change log with tokens; push is a hint, poll is the guarantee; 30-day retention, then resync." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| Users / daily actives (design assumption) | 500M / 150M; ~60M enterprise seats | Sets every rate below |
| Event writes (create, edit, RSVP) | ~1B/day; ~12K/s avg, ~60K/s Monday-morning peak | Write path is modest; fan-out behind it is not |
| Calendar view / agenda reads | ~2B/day; ~100K/s peak | Each read expands rules for a 1–5 week window |
| Sync "anything changed?" checks | ~400M connected clients ÷ 15 min backstop ≈ 450K/s | The single largest request class; must be a cheap cursor compare |
| Free/busy queries | ~1.5B/day; ~80K/s peak; ~6 calendars × 2 weeks each | Why an index beats expanding rules per query |
| RRULE expansion cost | ~1–5 µs per instance in a compiled engine | A year of a daily series is ~2 ms: fine per read, not per free/busy probe |
| Busy-index horizon | 18 months materialized; beyond that, expand on read | Covers nearly all scheduling; keeps the index ~70 KB per active calendar |
| Reminder rate | ~1.2B/day server-side; minute at :50 ≈ 20–30× average | The top-of-hour herd: most meetings start on :00 or :30 |
| Reminder lateness target | p99 ≤ 30 s | Users notice a 10-minute reminder arriving at 8 minutes |
| tz database releases | 2–7 per year (seven in 2022, two in 2024) (tz NEWS) | Rebase must be routine, not an incident |
| Shortest tz notice seen | ~2 days (Chihuahua, tz 2022f) (tz NEWS) | Rebase SLO must be hours, end to end |
| Change-log retention for sync | 30 days (design choice) | Longer-offline clients pay a full resync; shorter retention means resync storms |
| Free/busy staleness SLO | ≤ 5 s from write to index | Beyond that, people double-book rooms they "saw" free |
| Google free/busy request bounds | calendarExpansionMax ≤ 50, groupExpansionMax ≤ 100 (freebusy.query) | Bound fan-out per query; groups are not a free lunch |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on the month-view UI, an
eventstable and a REST API, then meet the first hard question — "what happens to this weekly meeting when the organizer's country drops daylight saving?" — with no answer, because they already threw away the zone. Compress the CRUD to ~8 minutes and spend the rest on the time model, recurrence, the invite ownership split, free/busy and reminders.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Users create single and recurring events, invite people inside and outside the company, and those people RSVP. Anyone scheduling a meeting can see when attendees and rooms are busy without seeing what they're doing. Reminders fire before events by push, email or on the device. Phones, web and third-party clients stay in sync, including offline. Hosts can publish a booking page where others pick a free slot."
Then the non-functional requirements, which is where the design lives:
"Four constraints drive everything. One: time correctness — an event shows the same local time on every device, survives DST, and survives governments changing the rules; that decides the storage model. Two: recurrence — series are infinite, so range queries must work without infinite rows. Three: free/busy for a meeting with a dozen people and a room must come back in under 200 ms and be fresh within about 5 seconds, or people double-book. Four: reminders must fire within 30 seconds of the right moment and must follow every edit. Scale: 500 million users, 150 million daily, 60 million enterprise seats, about a billion event writes a day, and roughly 400 million connected clients syncing."
Then name the underspecified parts:
"A few things I'd confirm: is this consumer, enterprise, or both? Do we need to interoperate with Exchange and CalDAV clients, or only our own apps? Are rooms and other resources in scope? Are booking pages with payments in scope? I'll assume both consumer and enterprise on one core, CalDAV and email interop, rooms in scope, and booking pages as a layer on top."
🎯 Staff Move: Saying "for future events, the source of truth is the local time and the zone; every UTC instant is a cache" in the first three minutes tells the interviewer you have already decided the hardest correctness question. Everything you draw next can be judged against it.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Calendar:
calendar_id,owner(user, group, room, resource),default_tz(IANA ID),acl[](free/busy only, read details, write, delegate) - Event (series or single):
event_id,ical_uid,calendar_id(organizer's),start_local,end_localorduration,tzid,all_day,rrule,rdates[],exdates[],sequence,title,location,visibility,status - Exception (instance override):
event_id,original_start(RECURRENCE-ID), overridden fields only,cancelled - Attendance (overlay):
event_id,attendee(user, room or external email),partstat(needs-action, accepted, declined, tentative),role,personal_reminders[],hidden,last_seen_sequence - Busy interval (derived):
calendar_id,start_utc,end_utc,event_id,original_start,transparency,tzdb_version - Reminder timer (derived):
fire_at_utc,user_id,event_id,original_start,offset,channel,sequence - Change log entry:
calendar_id,seq(monotonic per calendar),event_id,op
API:
POST /v1/calendars/{cal}/events { start_local, tzid, duration, rrule?, attendees[] }
PATCH /v1/events/{id}?scope=all|this|following&original_start=... (If-Match: sequence)
POST /v1/events/{id}/rsvp { partstat, original_start? } (overlay write only)
GET /v1/calendars/{cal}/events?from=...&to=...&expand=true (range read, expanded)
POST /v1/freebusy { calendars[≤50], from, to } → { cal: [ {start_utc, end_utc}, ... ] }
POST /v1/rooms/{room}/book { event_id, original_start, idempotency_key }
GET /v1/calendars/{cal}/changes?sync_token=... → { changes[], next_sync_token } | 410
The scope parameter on PATCH is the most important field on this page: "this instance", "this and following" and "all" are three different writes with three different blast radii.
🎯 Staff Move: "An RSVP is a write to the attendee's overlay row, never to the event. That one choice means the organizer editing the title and 300 attendees answering at the same time never contend on the same row — and it's also exactly how iTIP separates REQUEST from REPLY."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one edit in 90 seconds:
- Organizer in London creates "Design review, Tuesdays 09:30,
Europe/London,FREQ=WEEKLY;BYDAY=TU" with 12 attendees and Room 4B. The event service writes one series row, 13 overlay rows, and one change-log entry per affected calendar — in one transaction on the organizer's shard, with overlays for other shards written through the change stream. - The room is booked by a conditional insert of the first 18 months of instances into the room's busy set; any overlap rejects that instance and the organizer sees which dates conflict.
- Projectors expand the rule for 18 months, convert each instance to UTC with the current tzdb, and write busy intervals for every attendee who hasn't declined.
- The reminder projector writes timers for instances in the next 48 hours; a sweeper extends the horizon hourly.
- Attendees' devices get a push hint, call
changes?sync_token=…, receive the series and their overlay, expand locally, and schedule local alarms. - External attendees get an iMIP REQUEST with UID and SEQUENCE 0; their REPLY comes back through the iMIP gateway and updates only their overlay row.
- A New York attendee sees 04:30 in winter and 05:30 for the two or three weeks in March (and the one week around early November) when the US and UK are out of step on DST — because London is the anchor, which is what the organizer meant.
🎯 Staff Move: Say out loud: "Notice that only one box holds truth. The busy index, the reminders, the change logs and the outbound invites are all projections of the event store. If any of them is wrong — after a bug, a tz release, or a missed message — I can rebuild it from intent. That's the property I'll protect for the rest of the interview." You've now spent ~9 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes a calendar hard is that most of what it stores is in the future, and the future can change underneath it. I'd like to go deep on four things: how we represent time so DST and tz rule changes don't corrupt events; how recurrence and exceptions work when someone edits 'this and following'; how free/busy and room booking stay correct across a dozen people; and how reminders and device sync follow every change. Where would you like to start?"
If no preference: start with the time model. It's the question that decides the level.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: The time model (7–8 min)
"Three kinds of time, stored three ways. A future timed event is a local start, a duration, and an IANA zone ID — the organizer's zone unless they pick another. An all-day event is a date or date range, half-open, with no zone; Christmas is the 25th everywhere. A floating event — 'medication at 08:00' — has a local time and no zone, belongs only to its owner, and never enters a shared free/busy index. Once an instance is in the past, I freeze its UTC instant: an audit of 'when did this meeting happen' must not change because a tz release corrected history."
Then DST, precisely: "Two edge cases. Spring-forward: 02:30 on the transition day doesn't exist in New York. RFC 5545 says a recurring instance generated there is dropped; a single event at 02:30 is read with the pre-gap offset, landing at 03:30. I'd shift recurring instances forward too — users don't expect a meeting to vanish — and I'd document that as our policy and test it in every client. Fall-back: 01:30 happens twice; we take the first occurrence, as the RFC does."
Quantify: "A tz release changes a handful of zones. If a changed zone holds 5 million users with ~40 active series each, that's 200 million series to re-expand, but only future instances in the 18-month busy horizon — roughly 15 billion interval rows at worst, realistically far fewer because only instances after the effective date move. At 1 million rows a second across the projector fleet, about four hours. That's the SLO I'd publish: tz release to consistent index in under 6 hours."
Who pays: "Users in the affected zone pay a window where the server has new rules and their phone doesn't — or the reverse. I'd render server-computed times in our apps for affected zones until device tz data catches up, and show an in-app notice. Platform pays for the rebase pipeline and the version-skew dashboard."
Deep dive 2: Recurrence, exceptions and series edits (7–8 min)
"A series is a row with an RRULE plus RDATE and EXDATE lists. An exception is a row keyed by (event_id, original_start), holding only the overridden fields. Range reads expand the rule into the window, drop EXDATEs, add RDATEs, and overlay exceptions by original start. Three edit scopes: 'this instance' writes or updates one exception; 'all' rewrites the series and bumps SEQUENCE, keeping exceptions whose original start still matches the rule; 'this and following' splits the series — the old one gets UNTIL just before the split point, the new one starts there with a new ID linked by split_from, and exceptions after the split are re-homed to the new series."
Quantify: "Expansion is 1–5 µs per instance. A week view with 40 series is 40 instances — microseconds. The expensive case is a daily series with no end opened in a year view: 365 instances, about 1 ms. I cap expansion per request at 5,000 instances and paginate beyond it."
Who pays: "When the organizer changes a series' time, exceptions keyed to old original starts would be orphaned. I keep them if their original start is still generated by the new rule, and otherwise surface 'n instances had custom changes that were discarded' to the organizer before they confirm. The organizer pays a confirmation dialog; attendees don't pay with ghost meetings."
Deep dive 3: Free/busy and rooms (6–7 min)
"Free/busy reads a per-calendar index of busy intervals in UTC — instance-level, already expanded, already filtered for declined RSVPs and transparent events. A query for 12 people and 3 candidate rooms over two weeks is 15 range scans on 15 shards, merged into free windows in the requester's zone. Each scan returns maybe 40 intervals. The response contains intervals, never titles, unless the requester has delegate access."
"Rooms are different: a room is a resource with a hard invariant — no overlapping accepted bookings. Booking a room is a conditional insert into the room's interval set, enforced by the store — in PostgreSQL an exclusion constraint on (room_id, tstzrange) — not a free/busy check followed by a write. That's the Hotel & Home Booking answer applied to 30-minute slots."
Quantify: "80K free/busy queries a second at ~6 calendars each is ~500K range scans a second, each a few KB from a hot partition. The index is ~70 KB per active calendar for 18 months — about 35 TB across 500M calendars before replication, with the active working set far smaller."
Deep dive 4: Reminders and sync (5–6 min)
"Reminders are a projection: for every instance in the next 48 hours, for every attendee who hasn't declined, for every reminder they've configured, one timer keyed by (user, event, original_start, offset). Any edit, RSVP, or tz rebase re-derives timers for that event; the timer carries the event SEQUENCE so a stale timer that fires after an edit is dropped at delivery. Firing is a Job Scheduler problem; delivery is a Notifications problem. Devices also schedule local alarms from synced data, so a phone in airplane mode still rings; the server suppresses push for devices that confirm local scheduling."
"Sync is a per-calendar change log with a monotonic sequence. A client holds a token per calendar; 'anything changed?' is a compare against the latest sequence in a cache — 450K/s, almost all answered 'no' from memory. Push is a hint; the 15-minute poll is the guarantee. Tokens older than 30 days get a 410 and a full resync — paced, because a bug that invalidates tokens for 10 million clients is a self-inflicted outage."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: intent is the truth, instants are a cache. Future events are local time plus IANA zone plus rule; past events freeze in UTC. Series are rules with exceptions keyed by original start. One organizer-owned event with per-attendee overlays, iTIP at the boundary. Free/busy, reminders and sync logs are projections that a tz release, an edit or a bug can rebuild. And rooms get a hard no-overlap constraint, not a check-then-write."
The evolution closer:
"What I'd build later: a shared recurrence engine with a conformance suite across server, web and mobile; a tz version-skew dashboard; smart scheduling that suggests times using the same busy index; and booking pages with payment holds on the reservation path. What I'd not build: our own tz database or a custom invite protocol — iCalendar, iTIP and CalDAV already exist and every client speaks them."
🎯 Staff Move: End on the correctness failure you designed out and who owns it. Senior candidates end with "and we shard by user." Staff candidates end with "and
cal.tzdb.rebase_lag_hours,cal.freebusy.staleness_sandcal.reminder.lateness_p99are on the calendar platform's dashboard, because those three numbers tell us whether people are showing up at the right time."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| UI tour | Designs day/week/month views and drag-and-drop | One sentence: "Clients render ranges from the range-read API" |
| UTC by reflex | "All timestamps UTC" in Phase 1 | "UTC for the past; local + zone for the future" in Phase 1 |
| Row-per-instance | Materializes two years of every series | "Store the rule; materialize only busy intervals, 18 months" |
| Invites as copies | Fans every edit out to N calendars | "One event, overlays per attendee" in Phase 2 |
| Free/busy as a join | Expands rules for 12 calendars per query | Names the busy-interval index and its staleness SLO |
| Reminders as an afterthought | "And a cron job sends reminders" at minute 44 | Reminders as a projection with a dedupe key and a lateness SLO |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
A calendar is the cleanest example of a system whose data means something different tomorrow than it does today. Almost every other system stores facts: a payment happened, a message was sent, a file has these bytes. A calendar stores promises about the future in civil time, and civil time is defined by legislatures, not by physics. That makes the obvious engineering default — normalize everything to UTC — a correctness bug for exactly the data the product exists to hold.
It also has a deceptive happy path. A Senior engineer can build a calendar that works perfectly in a demo: a table, a REST API, a week view, invites. The design is judged entirely by days nobody demos: the weekend the US switches to DST and Europe hasn't, the Thursday a government announces it is dropping DST on Sunday, the morning the organizer drags a series with 40 exceptions to a new time, and the 08:50 when a third of the company's meetings fire reminders in the same second. Staff candidates design for those days before drawing the first box — and they notice that most of the system is derived views of a small amount of intent, which is what makes it repairable.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"An organizer in Phoenix creates a weekly 09:00 meeting with attendees in Denver and London. Arizona doesn't observe DST; Denver does; the UK does, on different dates. What time does each attendee see in the second week of March, the first week of April, and the first week of November — and what do you store so that a future change to any of those rules is handled correctly?"
A candidate who answers with the anchor (the meeting is 09:00 America/Phoenix, which is always UTC−7, so the instant never moves), the Denver attendee (09:00 MST in winter, 10:00 MDT once the US springs forward on the second Sunday of March, and back to 09:00 after the first Sunday of November), the London attendee (16:00 GMT in winter, 17:00 BST from the UK's spring-forward in late March), and the storage (local time, America/Phoenix, RRULE — and derived UTC instants tagged with the tzdb version, recomputed if Arizona's rules ever change) has built a calendar. A candidate who says "we store UTC and convert on display" has built a database of timestamps — and cannot tell you what the meeting was supposed to be.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Personal and shared calendars (consumer). The population is individuals and families: 500 million users, several devices each, intermittent connectivity, and a long tail of third-party clients. Most events are created on phones; many are recurring (birthdays, classes, routines); many are all-day or floating. The threats are correctness drift between devices, reminders that fire wrong or twice, and sync that loses an offline edit. The design centers on the intent-based time model, expand-on-read recurrence (so phones can expand locally too), a change-log sync protocol with tokens, and device-local alarms backed by server reminders. Correctness bar: every client shows the same local time for every instance; reminder lateness p99 ≤ 30 s; no lost offline edits.
Enterprise scheduling (free/busy, rooms, delegation). The population is companies: 60 million seats, executives with assistants who manage their calendars, thousands of rooms per campus, and meetings that cross time zones and company boundaries. The core query is "when are these 12 people and one room all free in the next two weeks?" The threats are double-booked rooms, free/busy that lags edits, private event details leaking through availability, and Exchange or CalDAV interop that silently drops updates. The design centers on the busy-interval index, organizer-owned events with attendee overlays, rooms as resources with a hard no-overlap invariant, an ACL model with "free/busy only" as a first-class permission, and iTIP for external attendees. Correctness bar: free/busy fresh within 5 s, zero double-booked rooms, no detail leaks.
Booking pages and appointment scheduling. The population is hosts (consultants, clinics, recruiters, sales teams) and the strangers who book them. A host publishes availability rules — working hours in their zone, buffers, minimum notice, maximum bookings per day — and a booker picks a slot shown in their zone. Traffic is bursty: a link shared in a newsletter produces thousands of availability reads in minutes for one host — a hot key. The design centers on computing availability from the same busy index, then committing a booking as an idempotent conditional insert that also blocks the host's calendar, with optional short holds while a payment completes. Correctness bar: zero double bookings, slot times correct across zones, booking confirmation in under 2 s.
🎯 Staff Move: "These share the time model, recurrence and the busy index, and almost nothing else. Consumer calendars need offline sync and local alarms; enterprise needs privacy-aware free/busy and room invariants; booking pages need a reservation-style write path under bursty load. I'll build one core and give each intent its own write and read policies on top."
2.2 When NOT to Build a Calendar Service#
| Situation | What to Do Instead | Why |
|---|---|---|
| Your product needs scheduling but calendars aren't the product | Integrate with users' existing calendars via their providers' APIs or CalDAV | Users won't move their calendar to you; you need their free/busy, not their events |
| Appointment booking for a small business | Buy a booking product that writes to the host's calendar | Recurrence, tz and interop are years of edge cases for a feature, not a business |
| System jobs that run "every night at 02:00" | A job scheduler — see Job Scheduler | Machines don't RSVP; a calendar's DST policy (shift, don't drop) is wrong for jobs that must not run twice |
| Hotel nights, rentals, inventory | A reservation system with date-based inventory — see Hotel & Home Booking | Nights are dates with counts and holds; calendar events have no inventory model |
| Shift rosters with labour rules | A workforce-management system that exports to calendars | Overtime, rest periods and union rules are constraints a calendar doesn't model |
| Reminders with no time-of-day semantics ("ping me in 2 hours") | A delayed message on a queue | Relative delays are UTC durations; no zone, no recurrence, no calendar |
| Public event listings (concerts, sports fixtures) | Publish an ICS feed or an "add to calendar" link | Readers subscribe; you never need their availability or RSVP |
And within the design, some things you should not build even when you own the service:
- Don't invent a recurrence format. Use RFC 5545 RRULE and its exception model; every client and every import/export already speaks it.
- Don't maintain your own tz data. Ingest IANA tz releases; your job is ingestion speed and skew measurement, not geopolitics.
- Don't put titles in the free/busy index. Keep it intervals only, so a bug in an ACL check cannot leak "Layoff planning" to the whole company.
- Don't treat rooms as people. A person can be double-booked and choose; a room cannot. Different invariant, different write path.
- Don't fire reminders from the event row. A timer created at event creation goes stale on the first edit; derive timers from current intent.
The Staff signal is knowing that a calendar is the right tool for human commitments in civil time and the wrong tool for machine schedules, inventory and relative delays — and that most calendar bugs come from mixing those three in one model. See Buy or Build: The Total-Cost Test and Drill 7.
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Say |
|---|---|---|
| Which zone anchors a meeting | Decides what moves when DST differs between attendees | "The organizer's zone by default; the organizer can pin another; attendees see it converted." |
| Interop scope | Exchange/CalDAV/iMIP adds round-trip and ordering constraints | "CalDAV and iMIP first-class; ICS subscribe read-only; Exchange via its API." |
| Free/busy freshness | Decides synchronous vs projected index | "≤ 5 s staleness; rooms checked synchronously at booking." |
| Rooms and resources | Hard invariant vs soft availability | "Rooms get a no-overlap constraint; people don't." |
| Max attendees | Fan-out of edits and reminders | "1,000 interactive guests; above that it's a broadcast event without per-guest RSVP tracking." |
| Offline duration | Change-log retention and resync cost | "30 days of change log; older clients do a paced full resync." |
| Privacy model | Who sees titles vs busy blocks | "Free/busy-only by default inside a company; delegates see details; nothing outside." |
2.4 Precise Terminology#
| Term | Meaning | Common Confusion |
|---|---|---|
| Instant | A point on the UTC timeline | Assumed to capture what the user meant |
| Wall-clock (local) time | Clock reading in a place, e.g. 09:30 | Treated as unique; it repeats in fall-back and skips in spring-forward |
| IANA zone ID | Europe/London, America/Chicago — a named rule history | Confused with abbreviations (CST is ambiguous) or fixed offsets (−06:00 has no DST) |
| UTC offset | Difference from UTC at an instant | Assumed constant for a zone |
| DST gap | Local times skipped at spring-forward (e.g. 02:00–02:59) | Events there "don't exist" and need a policy |
| DST overlap | Local times repeated at fall-back (e.g. 01:00–01:59 twice) | Ambiguous; RFC 5545 picks the first |
| Floating time | Local time with no zone | Shared by mistake; means a different instant for each viewer |
| All-day event | A date or date range, not an instant | Stored as midnight UTC, appears on the wrong day west of Greenwich |
| RRULE | Recurrence rule (RFC 5545) | Expanded once and thrown away |
| Original start / RECURRENCE-ID | The start an instance would have had from the rule | Confused with its current (moved) start |
| Exception | Override for one instance, keyed by original start | Orphaned when the rule changes |
| Series split | "This and following": old series ends, new one begins | Implemented as editing every future row |
| SEQUENCE | Organizer's revision number for an event (iTIP) | Incremented by attendees, or not at all |
| PARTSTAT | An attendee's participation status | Stored on the event, contending with organizer edits |
| Free/busy | Busy intervals without details | Confused with read access to the calendar |
| Transparency | Whether an event blocks time (OPAQUE) or not (TRANSPARENT) | All-day "working from home" blocks everyone's whole day |
3. Where the Design Splits#
Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.
3.1 Fault Line 1: Recurrence — Store the Rule vs Materialize Instances#
The tension: A rule is compact, editable in one write, and infinite. Rows are trivially indexable and queryable by range. Every calendar has to answer range questions over infinite rules; the only choice is where expansion happens and how much of it is kept.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Materialize every instance (N years) | Range queries are plain index scans; per-instance edits are row updates | Series "end" at the horizon; a time change rewrites N rows; exceptions become indistinguishable from generated rows; 40 attendees × daily × 2 years ≈ 29K rows per click | Organizers (lost exceptions); storage; every edit's write amplification |
| Store rule, expand on every read | One row per series; edits are one write; infinite series are free | Free/busy for 12 people expands 12 rule sets per query; every client needs an identical expander | Free/busy latency; client teams (N expanders) |
| Store rule + exceptions; expand on read; materialize busy intervals for a bounded horizon (Staff default) | Edits are one write; range reads expand cheaply; free/busy reads an index; the index is rebuildable | Two representations to keep consistent; horizon must roll forward; tz changes must rebase | Platform owns the projector and horizon sweeper |
| Client-only expansion (CalDAV-style "give me the master") | Server is simple | Server can't answer free/busy or fire reminders without expanding anyway | Every client; server features you can't build |
The subtle part — moved exceptions. An exception can move an instance into a window whose rule-generated instances don't include it (Tuesday's meeting moved to the following Monday), or out of the window its original start falls in. A correct range read therefore also queries exceptions whose current start overlaps the window, not only those whose original start does. Index exceptions by both.
Series edits, precisely:
| Edit Scope | Write | Exceptions | SEQUENCE |
|---|---|---|---|
| This instance | Upsert exception at original_start | Only this one | Incremented for the event (iTIP sends a REQUEST with RECURRENCE-ID) |
| All instances | Update series row | Kept if their original start is still generated; otherwise listed to the organizer and dropped on confirm | Incremented |
| This and following | Set UNTIL on old series to just before split; create new series with split_from; move exceptions after split to the new series | Re-homed | New series starts at 0; old series incremented |
RFC 5545 also allows RECURRENCE-ID;RANGE=THISANDFUTURE for "this and following". Inside the system I prefer the split, because it keeps every series' history immutable before the split point and makes each series a simple rule; at the CalDAV and iMIP boundary I translate either form.
The Staff default: store the rule, RDATE/EXDATE and exceptions keyed by original start; expand on read in a single shared engine; materialize busy intervals and reminders for bounded horizons (18 months and 48 hours) as projections.
When to deviate:
- Room and resource calendars: materialize instances into the room's booking table for the booking horizon, because the no-overlap constraint must be enforced per instance at write time (see 3.4).
- Very long horizons in reporting ("how many hours of meetings will we have next year?"): expand in a batch job, not the online index.
- Series with thousands of exceptions (a daily standup with years of moves): split the series automatically at a year boundary to bound per-read work.
3.2 Fault Line 2: Time Representation — UTC Instant vs Wall-Clock + Zone#
The tension: UTC instants compare, sort and index trivially and are the right default almost everywhere in backend engineering. But a future calendar event is a promise in civil time, and civil time's mapping to UTC is decided by governments and can change after you store it. Storing only the instant throws away the one thing you need to apply that change.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| UTC instant only | Simple; index-friendly | tz rule change → every future instance wrong, unrecoverably; DST drift for recurring series computed with old rules | Users in the affected zone; support; a migration that can't recover intent |
Local time + fixed UTC offset (09:00-06:00) | Looks precise | No DST: a winter offset applied in summer is wrong by an hour | Every user with a recurring event across a DST change |
| Local time + IANA zone, instants derived on read | Survives rule changes | No index: every free/busy query converts | Free/busy latency |
| Local time + IANA zone as truth; UTC instants derived into an index tagged with tzdb version; past frozen (Staff default) | Correct after rule changes; fast range queries; auditable past | A rebase pipeline; version skew between server and devices | Platform owns rebase and skew metrics |
DST, precisely. In the US, clocks spring forward from 02:00 to 03:00 on the second Sunday of March and fall back from 02:00 to 01:00 on the first Sunday of November; the EU and UK switch on the last Sundays of March and October. That gives a two-to-three-week window in March and a one-week window around the end of October when transatlantic meetings shift by an hour for one side only — correct behavior, and the most common "the calendar is broken" support ticket.
| Case | Example | RFC 5545 Rule | Our Policy |
|---|---|---|---|
| Single event in a gap | 02:30 New York on spring-forward day | Use the pre-gap offset → 03:30 EDT | Same, and warn at creation |
| Recurring instance in a gap | Daily 02:30 series | Instance is ignored and not counted | Shift to 03:30 (users don't expect a missing instance); documented and tested everywhere |
| Time in an overlap | 01:30 on fall-back day | First occurrence (EDT) | Same; offer "second occurrence" only via explicit UTC |
| Invalid date | Monthly on the 31st; yearly on 29 Feb | Ignored | Keep RFC default for BYMONTHDAY=31; offer "last day of month" (BYMONTHDAY=-1) in the UI; birthdays use RSCALE SKIP=FORWARD where clients support it |
tz data changes. Every derived row records tzdb_version. When a release arrives, the rebase job diffs old and new compiled rules per zone, finds zones whose offsets differ for any instant after now, and re-derives future instances in those zones only. Past instances — anything that ended before the release was ingested — keep their frozen UTC values. See Section 4.1 for the timeline when notice is two days.
Who signs off: product owns the gap/overlap policy and its wording in the UI; the calendar platform owns the rebase SLO and the tzdb ingestion; client teams own shipping the same policy. That's a written table, not a per-client judgment.
🎯 Staff Move: "I'm storing the meeting the way the organizer said it — 09:30, Europe/London — and treating every UTC instant as a cache tagged with the tz data version that produced it. The past I freeze. That one decision is why a government's two-day notice is a pipeline run for us and a data migration for everyone else."
3.3 Fault Line 3: Invitations — Shared Event vs Per-Attendee Copies#
The tension: Giving every attendee their own copy makes reads local and works naturally across systems — it is how email-based iTIP works between companies. But inside one system it turns every organizer edit into N writes that can partially fail, and makes the question "which version is true?" a distributed-consistency problem.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Full copy per attendee (fan-out on write) | Each calendar read is local; attendees can annotate freely | Edit = N writes; partial failure leaves stale copies; 1,000-guest event = 1,000 writes per typo fix | Attendees who see the old time; on-call reconciling copies |
| One shared event, no per-attendee state | One write per edit | RSVPs, personal reminders and "hide this" all contend on one row | Organizer (lock contention), attendees (no personal settings) |
| One organizer-owned event + per-attendee overlay rows (Staff default) | Edit = one write; RSVP = overlay write; no contention between them; one truth | Reads join event + overlay; cross-shard projections for attendees' indexes and logs | Platform owns projection; reads pay a join (cheap, same request) |
| iTIP copies at the boundary (external attendees) | Interoperable with every calendar system | Ordering, duplicates and lost REPLYs are normal | Integration team; SEQUENCE handling |
Ownership, precisely: the organizer's calendar owns start, end, tzid, rrule, title, location, attendee list and SEQUENCE. Each attendee owns partstat, personal reminders, colour and visibility of their instance. An attendee "moving" a meeting in their own calendar is either a COUNTER proposal to the organizer (iTIP) or a private copy — never a write to the shared event.
Large events. At 1,000+ guests, per-guest RSVP tracking and per-edit fan-out of notifications stop being useful and start being an attack surface. Above that threshold the event becomes a broadcast event: attendees subscribe, RSVP counts are aggregated, edits produce one change-log entry on the event's own feed rather than one per attendee calendar, and the organizer's row stops being a hot key for replies (see 4.3).
The Staff default: one event, organizer-owned, with per-attendee overlays inside the system; iTIP (REQUEST/REPLY/CANCEL with UID + SEQUENCE) at every boundary; broadcast mode above 1,000 guests.
When to deviate: cross-region tenants where an attendee's data must reside in another jurisdiction — store a replica of the event in the attendee's region and treat it as an iTIP-style copy with SEQUENCE ordering, accepting seconds of lag.
3.4 Fault Line 4: Free/Busy — Compute on Read vs Precomputed Busy Index#
The tension: Computing availability from events at query time is always consistent with the source of truth, but it expands every attendee's recurring series for every probe. A precomputed busy index answers in one scan per calendar, but it is a projection that can lag writes, miss a tz rebase, or disagree with the event store after a bug.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Expand events per query | Always exact | 12 calendars × 40 series × 2-week expansion per query; scheduling assistants issue a query per keystroke | Free/busy latency; event-store read load |
| Busy-interval index, synchronous update in the write transaction | Fresh immediately | Every edit writes N attendees' index partitions on other shards; distributed transaction per edit | Write latency and availability |
| Busy-interval index via async projection, ≤ 5 s lag (Staff default) | Fast reads; writes stay single-shard | Lag window; needs reconciliation | Users in the lag window; platform owns staleness SLO and reconciler |
| Index + synchronous check for hard resources | Rooms and booking pages correct at commit | Two paths to maintain | Booking service |
Rooms are the exception that proves the rule. A person double-booked is an inconvenience they resolve; a room double-booked is two meetings in the corridor. Room booking therefore does not trust the projected index. It is a conditional insert into the room's own instance table, with the invariant enforced by the store:
CREATE TABLE room_bookings (
room_id bigint,
during tstzrange, -- UTC, derived from local + zone
event_id bigint,
original_start timestamptz,
EXCLUDE USING gist (room_id WITH =, during WITH &&)
);
A recurring booking inserts each instance in the 18-month horizon in one transaction; conflicts come back per instance so the organizer sees "Room 4B is taken on 3 of 78 dates". The horizon sweeper extends room bookings as it extends busy intervals; a tz rebase of the room's zone re-derives them inside the same constraint. This is the Hotel & Home Booking inventory invariant applied to time ranges, and the same contention shape when 40 people want the one large room at 10:00 Monday.
Privacy. The index stores intervals, event_id and transparency — no title, location or attendees. Details are fetched only through the event service's ACL check. Free/busy permission is separate from read permission, and external requesters get coarser data (merged blocks, no event IDs).
The Staff default: async-projected busy index with a 5-second staleness SLO for people; synchronous conditional insert for rooms and booking pages; nightly reconciler comparing a sample of calendars' index against fresh expansion and reporting cal.freebusy.divergence.
When to deviate: small tenants (< 10K seats) can compute on read with a per-calendar cache and skip the index entirely; the index earns its complexity at scheduling-assistant query rates.
3.5 Fault Line 5: Reminders — Schedule at Write vs Rolling-Horizon Projection#
The tension: The simplest reminder is a timer created when the event is created. It goes wrong on the first edit, the first decline, the first series change, and the first tz release — and for an infinite series it can't be created at all. A projection that re-derives timers from current intent is always right but costs a re-derivation on every change.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Timer created at event creation | Simple | Stale after edits; can't represent infinite series; reminders for declined events | Users (wrong-time and ghost reminders) |
| Cron scans events every minute | Always current | Expands every user's rules every minute; 150M calendars | Infra cost; minute-boundary herd |
| Rolling 48 h projection, re-derived on change (Staff default) | Correct after edits, RSVPs and tz rebases; bounded timer count | Re-derivation fan-out on big series edits; sweeper must extend horizon | Platform owns projector and sweeper |
| Device-local alarms only | Works offline; zero server cost | Devices that haven't synced fire stale reminders; email and wearables need the server | Users with stale devices |
Dedupe and staleness. The timer key is (user_id, event_id, original_start, offset_minutes, channel); the timer carries the event SEQUENCE it was derived from. At fire time the worker compares against the event's current SEQUENCE: older means a re-derivation is in flight or missed, so it re-reads intent and either fires with fresh data or drops. That makes the timer store safe to be slightly behind — the check at delivery is the correctness backstop.
The top-of-hour herd. Most meetings start on :00 or :30, and most users keep the default reminder offset, so the minute at :50 carries 20–30× the average reminder rate. Jitter is not acceptable — a reminder that arrives two minutes late is a bug users notice. The answer is pre-staging: the timer service loads the next minute's timers into memory 60 s early, partitions them by user hash across workers, and hands batches to the notification system with a deliver_at; the notification system holds them until the exact second. Capacity is planned for the :50 peak, not the daily average. Timer firing mechanics are in Job Scheduler; per-channel delivery, collapse and quiet hours are in Notifications.
Device vs server. Phones schedule local alarms from synced data for the next 7 days and report "scheduled through sequence N" for each event; the server suppresses push for that device when the device's SEQUENCE matches. Email and wearables always use the server path.
The Staff default: rolling 48-hour timer projection re-derived from intent on every change, SEQUENCE check at delivery, pre-staged minute batches for the :50 herd, device-local alarms with server suppression.
When to deviate: all-day events and floating reminders (09:00 "on the day") are evaluated in the user's current zone, which the device knows better than the server — let the device own those and the server send only an email fallback.
4. When It Breaks#
4.1 Short-Notice Time Zone Change — Two Days to Move Every Future Meeting#
t=-2d 10:00: Government announces the region drops DST this weekend. tz maintainers
publish a new release the same day (the 2022f shape: two days' notice).
t=-2d 14:00: Without a pipeline: nothing happens until an engineer notices.
Server still on old tzdb. Busy index, reminders, room bookings
assume the old rules for 1.2M users in the zone.
t=-1d: Some phones receive an OS update with new tz data; most don't.
t=0 Sun: Clocks don't change locally. Server and stale phones think they did.
t=+1d 08:50: Reminders fire an hour wrong for recurring meetings in the zone.
Free/busy shows rooms free that are booked. Cross-zone attendees
see times that differ from what locals see.
t=+3d: Support volume 15× in the region. Engineers find no stored
intent for series written as UTC rows.
With the Staff design:
t=-2d 11:00: tzdb release detected by the ingestion watcher; compiled; diff shows
one zone changes offsets from Sunday 02:00 onward.
t=-2d 11:30: Rebase job: 1.2M users, 31M future instances in the 18-month horizon
re-derived; reminders and room bookings re-derived in the same pass.
t=-2d 13:00: Index tagged with new tzdb version. Clients in the zone are told to
render server-computed times until their device tzdb matches.
t=+1d: Meetings, rooms and reminders correct. Skew dashboard tracks devices
still on the old tzdb; in-app notice for those users.
Why it's hard: notice can be days, the server and every device carry independent copies of tz data, and the change must apply to future instances only. Lebanon in March 2023 is the extreme: a release moved its DST start, and another release five days later reverted it (tz NEWS). The rebase must be idempotent and reversible because the rules can flip back.
Detection: cal.tzdb.release_ingest_lag_hours (upstream release to server ingestion), cal.tzdb.rebase_lag_hours, cal.tzdb.version_skew (share of active clients on an older tzdb per zone), cal.render.server_client_mismatch (client reports of instances whose computed time differs from the server's).
Prevention: automated ingestion within an hour of release; rebase scoped by zone diff; every derived row tagged with its tzdb version; a "render server time" capability in clients; a runbook that includes customer comms.
Owner: calendar platform team (ingestion, rebase); client teams (render fallback); support (comms).
4.2 The :50 Reminder Herd#
t=08:49:00: Normal reminder rate ~15K/s for the region.
t=08:50:00: Default 10-minute reminders for every 09:00 meeting become due:
~400K timers due within the same second, 25× average.
t=08:50:02: Timer workers read due timers from the store by time bucket; the
08:50 bucket is one hot partition. p99 read latency 4 s.
t=08:50:30: Notification fan-out queue backs up behind the burst.
t=08:52:40: Reminders for 09:00 meetings still arriving. Users join late;
complaints "reminders are broken on Mondays".
The Staff design: time buckets are sub-partitioned by user hash (64 sub-buckets per minute) so one minute is not one partition; workers pre-load the next minute's timers 60 s early; batches go to the notification system with deliver_at and are held to the second; capacity is provisioned for the 99th-percentile minute of the week (Monday 08:50 local in the largest regions), not the average. The same hot-bucket problem appears in any time-indexed queue — see Hot Keys and Job Scheduler.
Detection: cal.reminder.lateness_p99 by minute-of-hour, cal.timer.bucket_read_ms, cal.reminder.due_backlog.
Owner: calendar platform (timers); notification platform (delivery capacity at the burst).
4.3 RSVP Storm on a Company-Wide Event#
t=0: Comms sends a 60,000-person all-hands invite from one organizer
calendar. Event treated as a normal event: 60K overlay rows, 60K busy
index writes, 60K change-log entries, 60K push hints.
t=+2min: 18,000 RSVPs in two minutes. Each REPLY also notifies the organizer
("Alex accepted"), writing to the organizer's change log.
t=+3min: Organizer's calendar shard: change-log append is a hot key at 150/s;
the organizer's own phone receives a sync hint per reply and resyncs
continuously. Shard p99 for every other calendar on it: 2 s.
t=+10min: Comms edits the dial-in link. Another 60K-way fan-out + 60K iMIP
REQUESTs to external partners on the list. Mail gateway throttled.
The Staff design: above 1,000 guests the event becomes a broadcast event: RSVPs aggregate into counters (sharded, like a leaderboard counter), the organizer gets a summary rather than per-reply notifications, attendee calendars reference the event through a subscription rather than per-attendee change-log entries, and edits publish once to the event's feed. External recipients above a threshold get one iMIP message per edit batched over 5 minutes.
Detection: cal.shard.hot_calendar_writes_per_s, cal.event.attendee_count at creation (warn above 1,000), cal.imip.outbound_queue_depth.
Owner: calendar platform; the internal comms tool owner signs off on broadcast mode as the default for distribution-list invites.
4.4 Sync Token Invalidation Storm#
t=0: A change-log compaction job is deployed with a bug: it truncates logs
to 2 days instead of 30.
t=+6h: Clients offline > 2 days present tokens older than the log: 410.
Each does a full resync: list every event in every calendar it holds.
t=+8h: Monday morning: 22M clients reconnect after the weekend; 9M get 410.
Full resync averages 1,800 events × 3 calendars per client.
t=+8h10m: Range-read tier at 6× normal load; event store read replicas saturated;
ordinary week-view loads time out. Clients retry full resyncs.
The Staff design: full resync is a planned, paced operation — the server returns 410 with a Retry-After drawn from a token bucket so resyncs are spread over hours; resync fetches a bounded window first (−30 days to +18 months) and backfills history lazily; compaction jobs assert a minimum retention and are canaried on one shard. This mirrors the cursor-reset failures in Cloud File Sync; Google documents the same 410 → full sync contract for its API (Google sync guide), and RFC 6578 defines it for WebDAV sync.
Detection: cal.sync.full_resyncs_per_min, cal.changelog.min_retention_days by shard, ratio of 410s to 200s on the changes endpoint.
Owner: calendar platform (sync); SRE for the canary policy on data-lifecycle jobs.
4.5 Silent Free/Busy Divergence#
A projector bug drops busy intervals for series whose RRULE contains BYSETPOS ("last weekday of the month") after a library upgrade. Nothing errors. Over three weeks, the scheduling assistant suggests slots that collide with monthly reviews; people decline, re-schedule, and blame the assistant. Rooms are unaffected (they use the synchronous path). Detection comes only from the nightly reconciler: it samples 100K calendars, expands from intent, diffs against the index, and reports cal.freebusy.divergence by RRULE feature — the spike is localized to BYSETPOS within one run. Prevention: the recurrence engine's conformance suite runs before any library upgrade, including BYSETPOS, negative BYMONTHDAY, DST gaps and invalid dates; the reconciler's divergence is a paging alert above 0.1%. Owner: calendar platform.
4.6 Room Double-Booked Through a Side Door#
The room booking path enforces the no-overlap constraint, but a facilities tool imports a weekly class schedule straight into room calendars as ordinary events, bypassing the booking service. Forty rooms are double-booked for a term. The fix is the same as in Hotel & Home Booking: every write path to a resource calendar — UI, API, CalDAV, iMIP auto-accept, admin imports — goes through the booking service; resource calendars reject direct event writes. Detection: cal.room.overlap_count from a periodic constraint audit; owner: calendar platform plus the facilities tool owner.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Short-notice tz change | cal.tzdb.rebase_lag_hours; cal.tzdb.version_skew | Every future event in the zone, plus cross-zone attendees | Automated ingest + scoped rebase; server-time rendering fallback | Calendar platform + client teams |
| Reminder herd at :50 | cal.reminder.lateness_p99 by minute-of-hour | Every 09:00 meeting in a region | Sub-partitioned buckets; pre-staging; deliver_at | Calendar + notification platforms |
| RSVP storm / hot organizer | cal.shard.hot_calendar_writes_per_s | Every calendar on the organizer's shard | Broadcast mode above 1,000 guests; aggregated RSVPs | Calendar platform |
| Sync token storm | cal.sync.full_resyncs_per_min | Range-read tier; all users' week views | Paced 410s; windowed resync; retention guards | Calendar platform |
| Free/busy divergence | cal.freebusy.divergence (reconciler) | Scheduling suggestions; trust | Conformance suite; nightly reconciler | Calendar platform |
| Room side-door writes | cal.room.overlap_count | Rooms on a campus | Single write path to resources | Calendar platform + facilities |
| Orphaned exceptions after series edit | cal.series.exceptions_dropped | One series' attendees | Confirm dialog; keep matching exceptions | Calendar platform + client teams |
| External iTIP out of order | cal.imip.stale_sequence_replies | One event with external guests | SEQUENCE ordering; REFRESH | Interop team |
🎯 Staff Insight: The dangerous calendar failures don't page. A meeting that is an hour off returns 200. A busy interval that is missing shows a free slot. A reminder for a declined meeting is delivered successfully.
cal.tzdb.version_skew,cal.freebusy.divergenceandcal.reminder.stale_sequence_dropsare the three metrics that turn silent correctness failures into tickets — make them first-class from day one.
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | Lists features: events, invites, reminders, views | Names consumer / enterprise / booking intents; commits; states "intent is truth, instants are cache" | Asks how many time libraries and tz data copies the company ships, and how skew is measured |
| Time | UTC everywhere, convert on display | Local + IANA zone for future, UTC frozen for past, floating and all-day handled; DST gap/overlap policy stated; tz rebase pipeline | tz data as a platform dependency with an ingestion SLO and a short-notice runbook across server and clients |
| Recurrence | Rows per instance, or RRULE with ad hoc edits | Rule + exceptions keyed by original start; three edit scopes; bounded materialization | One recurrence engine with a conformance suite as an org-wide contract |
| Invites | Copies per attendee | Organizer-owned event + overlays; iTIP with SEQUENCE at the boundary; broadcast mode for large events | Interop portfolio: which protocols are first-class, who funds them, and when to retire one |
| Free/busy & rooms | Query each calendar; rooms as attendees | Busy-interval index with staleness SLO; rooms as conditional inserts with a no-overlap constraint | Availability as a product API others build on, with privacy guarantees written down |
| Operations | "Add monitoring" | cal.tzdb.rebase_lag_hours, cal.freebusy.divergence, cal.reminder.lateness_p99, paced resyncs | Game days for tz changes and the Monday-morning herd; correctness SLOs reported alongside availability |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Separates intent from instant | "For future events the truth is 09:30 Europe/London; the UTC instant is a cache tagged with the tzdb version." |
| Knows DST edge cases precisely | "02:30 on spring-forward day doesn't exist; RFC 5545 drops recurring instances there. I'd shift them and document it." |
| Keys exceptions correctly | "Exceptions are keyed by original start, so moving the series doesn't lose the instance someone already moved." |
| Splits ownership | "An RSVP writes the attendee's overlay, never the event." |
| Treats derived state as rebuildable | "Busy index, reminders and change logs are projections; I can rebuild any of them from intent." |
| Names the real hot spots | "The biggest request class is 'anything changed?' and the biggest spike is reminders at :50." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Store all times in UTC" with no qualification | Cannot survive tz rule changes; loses intent for every future event |
| Materializes recurring events indefinitely or for a fixed N years | Series silently end; edits rewrite thousands of rows; exceptions lost |
Stores fixed offsets (-05:00) as the zone | Wrong by an hour across every DST transition |
| Rooms booked by "check free/busy, then insert" | Race between two bookers; double-booked rooms |
| Reminders scheduled once at creation | Fire for moved, declined and cancelled events |
| No answer for an offline client returning after weeks | Sync without a cursor-expiry story loses edits or melts the server |
5.4 Common False Positives#
- Fluency in date libraries ≠ a time model. Knowing
ZonedDateTimeis good; storing its output as UTC for a future meeting is the bug. - "We support RRULE" ≠ recurrence design. The questions are exceptions, edit scopes and where expansion happens.
- Knowing CalDAV exists ≠ sync design. The question is what happens when the token is older than the log.
- "Rooms are just attendees" sounds elegant ≠ correct. People can be double-booked by choice; rooms cannot.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Pick intents; "intent is truth, instants are cache"; scale |
| Entities & API | 3–5 min | Event, exception by original start, overlay, busy interval, timer, change log; edit scopes |
| Architecture | 5–10 min | ≤ 8 boxes; one source of truth; projections |
| Time model | 10–18 min | Local + zone, floating, all-day, frozen past; DST gap/overlap; tz rebase |
| Recurrence & edits | 18–25 min | Expand on read; exceptions; this/following/all |
| Free/busy & rooms | 25–32 min | Busy index, staleness SLO, privacy; room constraint |
| Pivot (interviewer's choice) | 32–42 min | Reminders, sync/CalDAV, booking pages, large events, multi-region |
| Wrap | 42–45 min | Intent vs cache; projections rebuildable; metrics and owners |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "A country drops DST with three days' notice" | Time model under change | Zone diff → scoped rebase → index, reminders, rooms; device skew fallback |
| "Find a slot for 15 people and a room" | Multi-calendar availability | Busy index scans, interval merge, working hours in each attendee's zone, room constraint at commit |
| "Add booking pages like a scheduling link" | Availability as a product + contention | Rules in host zone; slot display in booker zone; idempotent conditional insert; hot-link caching |
| "Support Outlook and Apple Calendar users" | Interop | CalDAV + iMIP; round-trip iCalendar; SEQUENCE; polling clients |
| "Reminders are late on Monday mornings" | Time-bucketed hot spots | Sub-partitioned buckets; pre-staging; capacity for the :50 minute |
| "Go multi-region with data residency" | Placement of truth | Home region per calendar; cross-region invites as iTIP-style copies |
6.3 What to Deliberately Skip#
- The month-view UI. "Clients render ranges returned by the range-read API."
- Search over event titles. "A per-user search index fed by the change stream; out of scope."
- Video-conference links. "An integration that writes a field on the event."
- Every RRULE part. "FREQ, INTERVAL, BYDAY, BYMONTHDAY, BYSETPOS, COUNT, UNTIL — we use a standard engine."
- Admin consoles and retention policy UI. "Policy is data the event service enforces."
6.4 Follow-Up Questions to Expect#
- "A weekly 09:00 meeting in Chicago has attendees in Berlin. What does Berlin see across March, and why does it change twice?"
- "The organizer moves 'this and following' for a series where three later instances were already moved individually. What happens to them?"
- "How do you guarantee a room is never double-booked when two people book it at the same time?"
- "A user declines one instance of a series. Which systems change, and how fast?"
- "A phone has been offline for six weeks. What does it download, and what does that cost us?"
- "The tz database ships a fix that changes a zone's past offsets. Do you change past events?"
- "How do you show free/busy to someone in another company without leaking event titles?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design Google Calendar."
Staff Answer
"Before drawing — is this consumer calendars, enterprise scheduling with free/busy and rooms, booking pages, or all three? They share a core and differ in write paths. I'll build the shared core and go deepest on enterprise scheduling.
The constraints I'll commit to: future events are stored as local time plus IANA zone plus recurrence rule, and every UTC instant is a derived cache tagged with the tz data version — past events freeze in UTC. Series are rules with exceptions keyed by original start. One organizer-owned event with per-attendee overlays; iTIP at the boundary. Free/busy reads a busy-interval index fresh within 5 seconds; rooms are booked by a conditional insert. Reminders are a rolling 48-hour projection that follows every edit. Scale: 500 million users, 60 million enterprise seats, a billion writes a day, 400 million syncing clients. I'll go: entities → one source of truth plus projections → time model → recurrence and edits → free/busy and rooms → reminders and sync."
Why this is L6:
- Distinguishes intents and commits
- States the time model before any boxes
- Frames free/busy, reminders and sync as rebuildable projections
What L7 adds:
- Asks how many recurrence engines and tz data copies already exist across server and clients
- Frames tz data as a company-wide dependency with an ingestion SLO
- Asks which interop protocols the business has promised enterprise customers
❌ Common L5 Trap
"An events table with start_time and end_time in UTC, a user_id, and an attendees table. Recurring events get expanded into rows for the next year. Reminders are a cron job that queries events starting in the next 10 minutes. Shard by user_id."
Why this misses: Each part works in a demo. The design cannot survive a tz rule change, ends every series after a year, rewrites hundreds of rows per series edit, expands nothing for free/busy, and scans every calendar every minute for reminders.
Drill 2: Core Mechanic — Expanding a Series With Exceptions#
Prompt: "Walk me through rendering next week for a user with a weekly series, one moved instance and one cancelled instance."
Staff Answer
"Load the user's single events whose start overlaps the week, series whose span overlaps it, and exceptions whose original or current start overlaps it. For each series: expand the RRULE in the series' zone between the window bounds, producing local start times; apply the DST policy (gap → shift forward, overlap → first occurrence) and skip invalid dates; convert each to UTC with the current tzdb. Remove instances whose original start is in EXDATE — that's the cancelled one. Overlay exceptions by original start — the moved one takes its new start, which may now fall outside the week, and an exception moved into the week from outside is added because I queried exceptions by current start too. Join the user's overlay for PARTSTAT and personal reminders. Convert to the viewer's zone for display. Cost: tens of instances, microseconds each."
Why this is L6:
- Expands in the series' zone, not the viewer's
- Handles exceptions by original start and by current start
- Applies a stated DST policy during expansion
What L7 adds:
- Requires every client to use the same engine or pass the same conformance suite
- Measures cross-client disagreement in production (
cal.render.server_client_mismatch)
Drill 3: Make It Concrete — Capacity#
Prompt: "Size it. Storage, reads, the busy index, reminders."
Staff Answer
"Writes: ~1B a day — creates, edits, RSVPs — ~12K/s average, ~60K/s on Monday mornings. Each fans out to projections: on average ~4 attendees, so ~50K busy-index writes a second at average and a few hundred thousand at peak.
Event store: ~1.5 KB per event row, ~120 bytes per overlay; ~200 TB a year of new event data before replication. Shard by calendar_id; an organizer's events and their overlays co-locate on the organizer's shard.
Busy index: an active calendar has ~3,000 instances in 18 months at ~24 bytes each ≈ 70 KB; 500M calendars ≈ 35 TB, most of it cold. Free/busy at 80K queries/s × ~6 calendars = ~500K range scans/s.
Sync: 400M clients; push hints plus a 15-minute backstop poll = ~450K/s 'anything changed?' checks, answered from a cached latest-sequence per calendar — a few hundred bytes in memory per active calendar.
Reminders: ~1.2B/day, ~14K/s average, ~300–400K/s in the minute at :50 in the largest region. 48-hour timer store ≈ 2.4B timers × ~100 bytes ≈ 240 GB."
Why this is L6:
- Finds the dominant request class (sync checks) and the dominant spike (:50)
- Sizes derived state separately from truth
- Ties shard key to ownership (organizer's calendar)
What L7 adds:
- Prices the device-vs-server split: every reminder a phone fires locally is server capacity not bought for the :50 peak
- Sets the change-log retention by cost of resyncs vs storage, not by convention
Drill 4: The Dependency Goes Down#
Prompt: "The projector pipeline is down for 40 minutes. What breaks?"
Staff Answer
"Writes still commit — the event store and change log are the truth, and they're on the write path; projections aren't. What goes stale: the busy index (new meetings don't block time), attendee change logs for cross-shard attendees (they don't see new invites), reminders for newly created or moved events within 48 hours, and outbound iMIP. Mitigations in order: free/busy marks calendars with projection lag over 30 s as 'possibly stale' in the UI; room bookings are unaffected because they're synchronous; reminder delivery re-reads intent at fire time, so a moved meeting doesn't fire at the old time — at worst a new meeting's reminder is late, so the timer service also does a just-in-time scan of events created in the last hour when the projector lag alarm is active. When the pipeline resumes it replays from its Kafka offset; projections are idempotent upserts keyed by (calendar, event, original_start)."
Why this is L6:
- Distinguishes truth from projections and what each failure costs
- Uses the fire-time SEQUENCE check as a correctness backstop
- Idempotent replay as the recovery mechanism
What L7 adds:
- Sets a projection-lag SLO and an error budget, and ties projector deploys to it
- Runs a quarterly game day that pauses projection on one shard during business hours
Drill 5: Hot Key — The Booking Link That Went Viral#
Prompt: "A founder posts their booking link to 2 million followers. What happens?"
Staff Answer
"Two load shapes on one calendar. Reads: thousands of availability requests a second for one host. Availability is a pure function of (host rules, busy index, now) that changes only on bookings, so I cache the computed slot list per host per day with a version from the host's change-log sequence — every booking bumps it. The edge serves it; origin computes it once per change. Writes: hundreds of people trying to book the same 20 slots. Each booking is a conditional insert on the host's slot range — same no-overlap constraint as rooms — with an idempotency key from the booker's session, so retries don't double-book and losers get 'slot taken, here are the next three'. I'd add a per-host booking rate limit and a waitlist option. The host's calendar shard sees ~20 successful writes and a few hundred rejected ones; the cache absorbs the rest. This is the Ticket Drops shape at small scale."
Why this is L6:
- Separates read hot key (cacheable, versioned) from write contention (conditional insert)
- Idempotency for retries
- Bounds damage with a per-host limit
What L7 adds:
- Offers hosts "high-demand mode" (lottery or queue) as a product choice, not an engineering patch
- Notes the abuse vector: bots booking out a competitor's calendar, and who owns that policy
Drill 6: Multi-Tenant — Free/Busy Across Companies#
Prompt: "Company A wants to see Company B's free/busy for scheduling. How?"
Staff Answer
"Within one provider: a tenant-level sharing policy, set by B's admin, that grants A's users free/busy-only access to B's calendars — merged blocks, no event IDs, no titles, horizon capped at, say, 60 days, rate-limited per requesting tenant. Across providers: the standards path is a CalDAV free-busy-query or a VFREEBUSY request via the scheduling outbox; in practice it's usually provider-to-provider federation configured by both admins. The design property is the same: the busy index stores no details, so a mis-scoped grant leaks availability, never content. Every cross-tenant query is logged for B's admin."
Why this is L6:
- Free/busy permission separate from read permission
- Coarser data across trust boundaries
- Auditable to the data owner
What L7 adds:
- Treats cross-tenant availability as a contract with legal review and a default-off posture
- Decides which federation protocols the company commits to supporting for years
Drill 7: Build vs Buy#
Prompt: "We're a recruiting product. Should we build our own calendar?"
Staff Answer
"No. Our users' calendars live in their existing providers, and they won't move. What we need is their free/busy and the ability to write interview events into their calendars. So: integrate with the major providers' APIs plus CalDAV, read free/busy, write events with our own UID namespace, and receive changes via push hints plus sync tokens — treating the 410 'token expired' path as normal. We build what is ours: interview-loop scheduling logic, panel constraints, candidate booking pages. We store a copy of availability with a short TTL and never become the system of record for anyone's time. Build cost for our own calendar would be a team of 8–10 for years — recurrence, tz, interop — to re-solve problems that aren't our product. See Buy or Build: The Total-Cost Test."
Why this is L6:
- Locates the system of record correctly (the user's provider)
- Names what to build (scheduling logic) vs integrate (calendar)
- Plans for sync-token expiry as a normal event
What L7 adds:
- Prices provider API quotas and the risk of one provider changing terms
- Builds an internal availability abstraction so providers can be added or dropped without touching scheduling logic
Drill 8: Changing Policy Without an Outage — Shift Instead of Drop#
Prompt: "Today recurring instances in a DST gap are dropped, per the RFC. Product wants them shifted forward. How do you roll that out?"
Staff Answer
"This changes the meaning of existing series, so it's a data-semantics migration, not a code tweak. Step one: version the expansion policy — gap_policy=omit|shift — and stamp it on every series; existing series default to omit. Step two: shadow — compute both expansions for series with instances in gaps over the next 18 months and report how many differ; it's tiny (series at 02:00–02:59 local in DST zones), maybe tens of thousands. Step three: ship the new policy in every client behind a capability flag; the server only flips a series to shift once its attendees' clients support it, or renders server-side for those that don't. Step four: flip new series to shift, then migrate old ones, re-deriving busy intervals and reminders for each. Step five: tell organizers whose series gain an instance. iCalendar export stays RFC-correct: we emit an explicit RDATE for shifted instances so other systems see them too."
Why this is L6:
- Treats semantic changes as migrations with a shadow phase
- Coordinates server and client rollout through capabilities
- Keeps interop correct by emitting explicit RDATEs
What L7 adds:
- Writes the policy into the org-wide time standard so every product that expands recurrences follows it
- Considers whether to propose the behavior upstream rather than diverge silently
Drill 9: Multi-Region and Data Residency#
Prompt: "EU customers require their calendar data to stay in the EU. Meetings span EU and US users."
Staff Answer
"Each calendar has a home region; truth for an event lives in the organizer's calendar's home region. An EU-organized meeting with US attendees: the event stays in the EU; US attendees' calendars hold a reference plus the minimum projection allowed by policy — times and busy status, and details only if the tenant's policy permits — replicated as an iTIP-style copy with SEQUENCE ordering. RSVPs from US attendees travel back to the EU as REPLY-style messages. Free/busy across regions reads each calendar's index in its home region; queries pay one cross-region hop (~80–120 ms transatlantic), which fits a 200 ms budget if fanned out in parallel. Reminders fire from the attendee's home region using the projected copy. During a region partition, each side keeps working on its own calendars; cross-region invites queue and apply in SEQUENCE order on heal. See Multi-Region."
Why this is L6:
- Places truth by ownership and residency, not by user location
- Reuses the iTIP ordering model internally across regions
- States the latency cost and the partition behavior
What L7 adds:
- Negotiates with legal what a "minimal projection" is, and makes it a tenant-visible setting
- Prices cross-region replication and decides which regions get full stacks vs read-through
Drill 10: Cost#
Prompt: "Calendar infrastructure costs are up 40% year over year. Where do you look?"
Staff Answer
"Three suspects, in order. Sync: 'anything changed?' checks scale with connected clients times poll frequency, not with users; one client version polling every 60 s instead of 15 minutes multiplies that line by 15. I'd break down request rate by client version and app. Resyncs: full resyncs cost thousands of events each; if cal.sync.full_resyncs_per_min is up, find what's expiring tokens. Projections: busy-index write amplification from large events and series edits — one 60K-person event edited ten times is 600K index writes. Fixes: enforce client poll floors at the gateway, broadcast mode above 1,000 guests, lazy materialization for calendars that haven't been queried in 90 days (expand on read for them; most consumer calendars are never free/busy-queried by anyone else)."
Why this is L6:
- Knows which load scales with clients rather than users
- Ties cost to specific amplification paths with metrics
- Proposes lazy materialization based on query patterns
What L7 adds:
- Moves work to devices deliberately (local expansion, local alarms) and measures server savings
- Sets per-partner API quotas so third-party clients carry their own cost
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — Monday 08:50, Reminders Two Minutes Late#
Context: For the third Monday running, users in the largest region report 09:00 meeting reminders arriving at 08:52 or later. The timer service is "healthy": no errors, CPU at 40%. The VP of the enterprise product escalates to you after a customer's CEO missed a board call.
Questions to Surface First:
- Is the lateness uniform, or concentrated in specific minutes of the hour?
- Where does the time go — reading due timers, handing off to notifications, or delivery by the push provider?
- Did anything change in timer partitioning or default reminder offsets recently?
Typical L5 Approach: Scales the timer worker fleet 3×. Lateness barely moves, because every worker is reading the same hot time-bucket partition.
Staff Approach: Breaks
cal.reminder.lateness_p99down by minute-of-hour and stage: lateness is only at :50 and :20, and it's all in the bucket read. The 08:50 bucket is one partition holding ~400K timers. Re-keys buckets as(minute, user_hash mod 64), adds pre-staging (load next minute 60 s early), and passesdeliver_atto the notification system so delivery is on the second.
Principal Approach: Treats the top-of-hour herd as a capacity class across every time-triggered system in the company — reminders, scheduled reports, cron jobs — and sets a shared rule: time-indexed stores must sub-partition buckets and plan for the 99th-percentile minute. Also asks product whether a 9-minute default reminder offset for some users would flatten the peak without hurting experience.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm lateness by minute-of-hour; identify the hot bucket partition |
| Triage | Timer bucket keyed by minute only; one partition per minute; read p99 4 s at :50 |
| Quick fix | Pre-load the next three peak minutes into worker memory ahead of time for this week |
| Guardrails | Sub-partitioned buckets; deliver_at hand-off; load test the :50 minute weekly |
| Post-mortem | Why was capacity planned on average rate? Why didn't the dashboard show lateness by minute-of-hour? |
Metrics to Watch: cal.reminder.lateness_p99 by minute-of-hour, cal.timer.bucket_read_ms, notif.deliver_at_hold_ms
Organizational Follow-up: reminder lateness becomes an SLO with an error budget owned by the calendar platform; notification platform commits to burst capacity at :50.
Ownership Question: "Who owns a late reminder — calendar or notifications?" Staff answer: Calendar owns 'due timer handed off before deliver_at'; notifications own 'delivered by deliver_at + 10 s'. Each half has its own SLO, and the end-to-end probe pages calendar first.
Key Takeaway: "Calendar load is shaped by the clock on the wall. Plan for the minute everyone shares, not the day's average."
What clears the Staff bar:
- Breaks lateness down by minute-of-hour and pipeline stage
- Fixes partitioning, not fleet size
- Splits the SLO between timer and delivery owners
Deep Dive 2: Silent Failure — Recurring Meetings an Hour Off in One Country#
Context: Support tickets trickle in from one country: some recurring meetings show an hour off, but only on some devices, and only for instances after last weekend. The country announced eleven days ago that it would not fall back this year. No alerts fired.
Questions to Surface First:
- When did the tz release with this change come out, and when did we ingest it?
- Which representation do affected series use — local + zone, or something else?
- Which clients disagree with the server, and which tzdb version do they carry?
Typical L5 Approach: Updates the server's tz library and restarts services. The server now renders correctly; the busy index, room bookings and reminders computed before the update are still on old rules, and phones on old OS tz data still disagree.
Staff Approach: Finds that ingestion is manual and lagged the release by nine days. Runs the rebase for the changed zone (future instances only), re-derives reminders and room bookings in the same pass, and enables server-rendered times for clients in that zone whose reported tzdb version is older than the fix. Then automates ingestion with a 1-hour SLO and adds
cal.tzdb.release_ingest_lag_hoursas a paging alert.
Principal Approach: Makes tz data a tracked dependency for the whole company — every service and client reports the tzdb version it runs; a dashboard shows skew by platform; and a short-notice runbook (ingest, rebase, client fallback, customer comms) is owned by the calendar platform and exercised twice a year with a synthetic zone change.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm the zone and effective date; check server tzdb version |
| Triage | Ingestion manual; server 9 days behind; devices split by OS update status |
| Quick fix | Ingest release; rebase the zone; server-render for stale clients |
| Guardrails | Automated ingestion; rebase lag alert; version-skew dashboard; reconciler checks a sample in each changed zone |
| Post-mortem | Why was a correctness dependency updated by hand? Why did no metric show server-vs-client disagreement? |
Metrics to Watch: cal.tzdb.release_ingest_lag_hours, cal.tzdb.rebase_lag_hours, cal.tzdb.version_skew by zone and platform, cal.render.server_client_mismatch
Organizational Follow-up: tz ingestion owned by the calendar platform with an SLO; client teams own reporting their tzdb version.
Ownership Question: "Who decides what users see when server and phone disagree?" Staff answer: The calendar platform, by policy: for zones with a pending change, the server's computation wins and clients render it, because the server can be updated in an hour and phones can take weeks.
Key Takeaway: "Time zone data is a production dependency with a release cadence. If nobody owns ingesting it, it is a scheduled outage."
What clears the Staff bar:
- Recognizes the fix spans index, reminders, rooms and devices — not just the server library
- Rebases future instances only
- Measures and mitigates version skew
Deep Dive 3: Large-Customer Onboarding — 400,000 Seats Migrating From Another Provider#
Context: A global enterprise is moving 400,000 users, 12,000 rooms and ~900 million events (including 15 years of history and ~40 million active recurring series) onto your platform over one weekend, with the old system still receiving edits until cut-over.
Questions to Surface First:
- What's the import format — iCalendar, a proprietary API, or both? How are exceptions and time zones represented?
- Which events are organized by people outside the migrating tenant?
- How do we handle edits in the old system during migration?
Typical L5 Approach: Bulk-imports events through the public API over the weekend. Rate limits stretch it to nine days; recurring series with custom (non-IANA) time zone definitions import as fixed offsets; room double-bookings appear where the old system allowed them.
Staff Approach: Builds a bulk import path that writes intent directly to the event store per shard and lets projections catch up afterwards (with free/busy marked "warming" for the tenant). Maps every source time zone definition to an IANA ID — Windows zone names and custom VTIMEZONE blocks go through a mapping table, and unmappable ones are flagged for review rather than silently converted to fixed offsets. Imports history as frozen UTC, future events as intent. Runs a delta sync from the old system until cut-over. Imports rooms with the no-overlap constraint in report mode first and hands facilities a conflict list before enforcing.
Principal Approach: Turns this into a repeatable migration product: a published import contract, a zone-mapping service, and a tenant-level "migration mode" that relaxes quotas and defers projections. Prices it into enterprise deals, because every large customer will need it and each bespoke migration costs a team for a month.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Separate bulk path from public API; decide history vs future handling |
| Triage | Inventory zone definitions; count series with exceptions; find room conflicts |
| Quick fix | Zone mapping table with manual review queue; rooms imported in report mode |
| Guardrails | Projection warm-up with a visible "warming" state; delta sync until cut-over |
| Post-mortem | Which mappings were ambiguous? How long did projections take per million series? |
Metrics to Watch: cal.import.events_per_s, cal.import.unmapped_tz_count, cal.projection.lag_s for the tenant, cal.room.overlap_count in report mode
Organizational Follow-up: migration tooling owned by a platform team, not the deal team; zone mapping maintained as shared data.
Ownership Question: "Who resolves the 3,000 room conflicts?" Staff answer: The customer's facilities team, with our conflict report, before enforcement is turned on. We don't silently pick winners for their rooms.
Key Takeaway: "A migration imports intent or it imports bugs. Fixed offsets are the bug."
What clears the Staff bar:
- Distinguishes history (frozen UTC) from future (intent) at import
- Refuses silent conversion of unknown zones to offsets
- Enforces invariants in report mode first
Deep Dive 4: Post-Mortem — A "This and Following" Edit Deleted Three Months of Customizations#
Context: An executive assistant changed the time of a weekly leadership meeting "this and following". Afterwards, 14 instances that had been individually moved or had different rooms reverted to the series defaults; two board-prep sessions lost their rooms. The assistant files a complaint; you lead the post-mortem.
Questions to Surface First:
- How does the split handle exceptions after the split point?
- Did the UI warn that custom instances would change?
- What happened to room bookings for re-homed instances?
Typical L5 Approach: Restores from backup for this one series and adds a confirmation dialog.
Staff Approach: Finds the split re-homed exceptions by original start, but the new series' rule generated different original starts (time changed from 10:00 to 11:00), so no exception matched and all were dropped. Fixes the split to re-key exceptions by mapping old original starts to new ones (same date, new time), preserving overridden fields; room bookings for re-homed instances are re-validated against the constraint and conflicts reported. Adds
cal.series.exceptions_droppedand a pre-commit diff shown to the user: "14 customized instances — keep customizations / reset them".
Principal Approach: Treats series edits as a class of destructive operation that needs an undo: every series edit records a reversible change set for 30 days. Adds series-edit scenarios (with exceptions, rooms and external attendees) to the recurrence conformance suite every client runs.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Restore the series and exceptions from the change log, not backup |
| Triage | Exception re-homing keyed on original start that no longer exists |
| Quick fix | Map exceptions by date across the split; preserve overrides |
| Guardrails | Pre-commit diff; cal.series.exceptions_dropped alert; 30-day undo |
| Post-mortem | Why could a single edit destroy data silently? Which other edit paths share this logic? |
Metrics to Watch: cal.series.exceptions_dropped, cal.series.split_count, undo usage rate
Organizational Follow-up: recurrence edit semantics documented and owned by the calendar platform; clients may not implement their own split logic.
Ownership Question: "Who owns the definition of 'this and following'?" Staff answer: The calendar platform. It's a server-side operation with one implementation; clients call it, they don't reimplement it.
Key Takeaway: "Exceptions are user data. Any edit that can drop them must show a diff and support undo."
What clears the Staff bar:
- Identifies the original-start keying as the root cause
- Restores from the change log, not a backup
- Adds a metric that makes silent data loss visible
Deep Dive 5: Multi-Region Expansion — Launching a Sovereign Region#
Context: The company is launching a region for public-sector customers who require that event data never leave the country. Their employees regularly meet with private-sector users in other regions.
Questions to Surface First:
- What may leave the region — nothing, busy blocks only, or times without titles?
- Who is the organizer of record for cross-region meetings?
- How do reminders and push notifications work if the push provider is outside the country?
Typical L5 Approach: Deploys a full stack in the new region and disables cross-region invites.
Staff Approach: Full stack in-region with home-region calendars. Cross-region meetings use the iTIP-style copy model: when a sovereign user organizes, the event stays in-region and outside attendees receive only what policy allows (time, organizer, a generic title); when an outside user invites a sovereign user, the invite crosses as an iMIP-like message into the region. Free/busy across the boundary returns merged busy blocks only, rate-limited. Reminders fire in-region; push hand-off uses a provider path approved for the region, and email reminders use an in-country relay.
Principal Approach: Writes the data classification for calendar fields once — time, busy status, title, description, attendees, location — and defines which may cross which boundary, so every future region is a configuration of that table rather than a new design.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Agree the field-level crossing policy with legal |
| Triage | Identify all cross-region flows: invites, free/busy, reminders, search, notifications |
| Quick fix | Cross-region iTIP-style copies with field filtering |
| Guardrails | Egress audit on every cross-region message; synthetic tests that titles never leave |
| Post-mortem | (Launch review) Which flows were discovered late? |
Metrics to Watch: cal.xregion.messages_per_s, cal.xregion.policy_violations (must be 0), cross-region free/busy p99
Organizational Follow-up: data classification table owned by privacy engineering; calendar platform enforces it.
Ownership Question: "Who approves a new field crossing the boundary?" Staff answer: Privacy engineering and legal approve; the calendar platform implements and audits. Product cannot add a field to cross-region payloads on its own.
Key Takeaway: "Cross-region calendars are federation, not replication — the same iTIP model you already use with other companies."
What clears the Staff bar:
- Reuses the iTIP copy model for cross-boundary meetings
- Filters fields by policy, not by convenience
- Covers reminders and notifications, not just storage
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why future events are stored as local time + IANA zone and past events as UTC, and what a tz release does to each
- State DST gap and overlap behavior precisely, including the RFC 5545 defaults and the policy you'd choose
- Model recurring series with exceptions keyed by original start, and implement this / this-and-following / all edits
- Separate organizer-owned event data from per-attendee overlays, and use iTIP semantics at the boundary
- Design a busy-interval index with a staleness SLO and privacy guarantees, and a no-overlap write path for rooms
- Treat reminders as a rolling projection with SEQUENCE checks and plan for the :50 herd
- Design change-log sync with tokens, push hints, poll backstops and paced full resyncs
The Bar for This Question#
Mid-level (L4): Builds events and attendees tables with UTC timestamps, a week view, and invites as copies. Recurring events are rows. Reminders are a cron query. Works in the demo; fails on the first DST change, series edit or offline device.
Senior (L5): Adds RRULE storage, attendees with RSVP status, a cache, sharding by user, and a reminder queue. Knows time zones exist and converts on display. The gap: UTC as the truth for future events, recurrence edits that lose exceptions, free/busy as a per-query expansion, rooms booked by check-then-write, and reminders created once. The design passes review and breaks on a government's announcement.
Staff+ (L6): Stores intent and treats instants as a cache, with a rebase pipeline tagged by tzdb version. Expands on read with exceptions keyed by original start. Splits organizer and attendee ownership. Answers free/busy from a projected index with a staleness SLO and protects rooms with a hard constraint. Makes reminders and sync re-derivable projections. Names who pays: users in changed zones during skew, organizers confirming destructive edits, the platform for rebase and reconciliation. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Store Everything in UTC" Is Wrong for Calendars#
| Claim | Reality |
|---|---|
| "UTC is unambiguous" | It is — and it's unambiguously wrong after a tz rule change for a future local-time commitment |
| "Convert on display" | Display can't recover the zone the organizer meant once it's been discarded |
| "Rule changes are rare" | tzdb shipped seven releases in 2022 and two-day notice is on record (tz NEWS) |
The Staff position: UTC is the right representation for the past and for machine timestamps. For future human commitments, the truth is local time plus zone, and UTC is a derived cache.
Why this matters in interviews: "UTC everywhere" is the most common correct-sounding wrong answer in this question. Qualifying it is an immediate level signal.
10.2 The RFC's DST-Gap Default Is Not What Users Want#
The Staff position: RFC 5545 says recurring instances at nonexistent local times are ignored (RFC 5545). A user with a 02:30 daily series doesn't expect a missing day. Shift forward, document it, test it in every client, and emit explicit RDATEs at export so other systems agree.
Why this matters in interviews: Knowing the standard and deciding to deviate deliberately is stronger than either alone.
10.3 Most Calendars Should Not Have a Materialized Busy Index#
The Staff position: The index earns its keep for calendars someone else queries — enterprise users, rooms, booking hosts. Most consumer calendars are never free/busy-queried by anyone else. Materialize lazily on first query and expire after 90 days without one; expansion on read is microseconds per instance.
Why this matters in interviews: Knowing where not to precompute shows cost judgment, not just correctness.
10.4 Rooms Are Inventory, Not Attendees#
The Staff position: Treating rooms as attendees is elegant and wrong. A person decides whether to accept a conflict; a room must not have one. Rooms get the reservation-system invariant and a single write path.
Why this matters in interviews: It shows you can tell when two things that look alike in the UI have different invariants underneath.
10.5 Push Notifications for Sync Are a Hint, Never a Delivery Guarantee#
The Staff position: Google's own documentation says its calendar push notifications are not 100% reliable and carry no body (Google push guide). Any sync design whose correctness depends on pushes arriving is broken; the poll with a cursor is the guarantee.
Why this matters in interviews: It separates candidates who have integrated with real calendar APIs from those who drew a WebSocket and stopped.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer designs a correct calendar. The Principal engineer notices that civil time is a company-wide dependency nobody owns. The calendar server has a tz library; the web client another; iOS and Android use the operating system's; the email digest renders times with a fourth; billing computes "end of month" in a fifth; the reminder and job schedulers each have their own DST policy. After every short-notice tz release these disagree for weeks, and the bug users report is never "wrong rule" — it is "my phone and my laptop show different times" or "my invoice is dated tomorrow". The L7 problem is not one calendar; it is one time model for the company: who ingests tz data and how fast, which recurrence engine everyone uses, which DST policy applies to people versus machines, and how skew is measured across every surface.
🧭 Principal Move: "Before I design the calendar, I want to know how many copies of the tz database we ship, how long it took each of them to pick up the last short-notice change, and how many recurrence expanders exist across our clients. The calendar's correctness is bounded by the slowest of those, so that's where I'd invest first."
The Org-Level Fault Line#
One time-and-recurrence platform vs each product owning its own.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product and client owns its time handling | Autonomy; fits each platform | N tz versions, N RRULE edge-case behaviors, N DST policies; skew after every release | Users (inconsistency); support; every post-mortem |
| One central time service everyone calls | One truth | Network call on every render; mobile offline breaks | Latency; offline users |
| Platform owns the contract, data and conformance suite; products embed a certified engine (Staff/L7 default) | Same behavior everywhere, works offline; tz updates shipped as data | Conformance suite and data pipeline must be maintained; clients must update | Platform (stewardship); client teams (adoption) |
🧭 Principal Move: "The platform owns what must be identical everywhere — tz data ingestion, the recurrence semantics, the DST policy for people versus machines, and the conformance suite. Products own rendering and UX. Any client that can't pass the suite renders server-computed times. That's how we make a two-day-notice change a data push instead of five incidents."
Cost Model#
Assumptions: fully loaded engineer ~$250K/year, cloud list prices, rough ranges (estimates, not quotes).
| Scale | Users / Seats | Infra ($/month) | Headcount | On-call Load | Buy Alternative |
|---|---|---|---|---|---|
| Startup feature (scheduling inside another product) | 100K users | ~$1–3K (integrations, cache) | 1–2 eng on provider integrations | Low; provider outages are the pages | Integrate with existing providers — always |
| Growth (own calendar for a vertical, e.g. clinics) | 5M users, 200K hosts | ~$30–80K (event store, index, timers, sync) | 8–12 eng: core, sync/interop, mobile, booking | Shared rotation; tz releases and Monday peaks are the pages | Partial: buy booking UI, own data |
| Planet-scale | 500M users, 60M seats | ~$3–8M (sync and reminder fan-out dominate) | 150–300 eng across core, interop, clients, rooms, enterprise admin | Dedicated rotations per tier; correctness SLOs | Not applicable |
The pricing insight: storage is never the cost driver. Connected clients times poll frequency, reminder peak capacity, and projection fan-out from large events are. Every reminder a phone fires locally and every week view it expands locally is server capacity you don't buy — so the client architecture is a cost decision.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Storing future events as local + zone vs UTC | One-way | Once intent is discarded it can't be recovered; migrating back means guessing zones for billions of events |
| Instance identity = (event, original start) | One-way | Every exception, RSVP, reminder and external client keys off it |
| Recurrence semantics (DST gap policy, invalid dates) | One-way-ish | Changing them moves existing instances for millions of series; needs a versioned migration (Drill 8) |
| iCalendar UID namespace for exported events | One-way | External systems keep UIDs forever; reuse or change creates duplicates |
| Organizer-owned event vs per-attendee copies | One-way-ish | Re-modelling invites touches every write path and every interop adapter |
| Busy-index horizon (18 months) | Two-way | A projection; rebuild with a new horizon |
| Event-store technology | Two-way | Behind the event service; dual-write and migrate per shard |
| Push vs poll mix for sync | Two-way | Client and gateway config |
🧭 Principal Insight: The time representation and instance identity are the decisions I'd slow down on. Stores, horizons and sync cadences can change in a quarter; "what does this event mean" and "which instance is this" cannot change at all once billions of events and every external client depend on them.
The Standard I'd Write#
RFC-TIME-001: Civil Time and Recurrence Standard
Status: Approved Owners: Calendar Platform + Client Platform Leads
Scope
Every service and client that stores, schedules, renders or bills against
human (civil) time: calendar, reminders, booking, notifications quiet hours,
billing periods, scheduled reports.
MUST
1. Store future human commitments as local time + IANA zone ID (or date, or
floating time), never as a UTC instant or fixed offset alone.
2. Store past events and machine timestamps as UTC instants.
3. Tag every derived UTC instant with the tzdb version used to compute it.
4. Ingest each IANA tz release within 1 hour of publication and complete
rebase of derived future instants within 6 hours.
5. Expand recurrences only with the certified engine, or a client that passes
the conformance suite for the current semantics version.
6. Apply the published DST policy: people-facing events shift out of gaps
and take the first occurrence in overlaps; machine jobs follow the job
scheduler's policy, documented separately.
7. Report the tzdb version in use from every client and service.
SHOULD
1. Render server-computed times when the client's tzdb is older than the
server's for the zone being displayed.
2. Emit explicit RDATE/EXDATE at iCalendar export wherever internal policy
differs from RFC 5545 defaults.
Exceptions
Filed with Calendar Platform; reviewed within 5 business days; time-boxed to
two quarters.
Success metrics
- cal.tzdb.release_ingest_lag_hours p99 ≤ 1
- cal.tzdb.rebase_lag_hours p99 ≤ 6
- cal.render.server_client_mismatch ≤ 0.01% of rendered instances
- Recurrence engines in production: 1 (plus certified ports)
What I'd Tell the VP#
"Calendars break on a schedule we don't control: governments change clock rules, sometimes with two days' notice, several times a year. Today, five parts of our product handle time separately, and each takes a different amount of time to catch up — so after every change, customers see different meeting times on their phone and their laptop. I'm proposing one owner for time handling, an automated pipeline that applies rule changes within hours, and a single recurrence engine every app uses. It's roughly six engineers for a year, mostly consolidating work we already do in five places. Success is measured by how fast we apply changes and how often our apps disagree, which I'll report quarterly. The cost is that client teams give up their own time code, and I'll make sure they get a library that's better than what they have."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Counts time implementations, not events | "How many copies of tz data and how many RRULE expanders do we ship?" |
| Prices the client architecture | "Every reminder the phone fires locally is :50 capacity we don't buy." |
| Sets the org's correctness posture | "Rebase within six hours of a tz release is an SLO with an owner and a game day." |
| Identifies the real one-way doors | "Local-plus-zone and instance identity are forever; the store isn't." |
| Knows when not to centralize | "Rendering and UX stay with products; semantics and data are the platform's." |
Staff answers that L7 interviewers find insufficient:
- "We store local time plus zone and rebase on tz changes" — correct for the calendar; silent on the other four surfaces that render or schedule the same events.
- "Clients use a standard RRULE library" — which one, at what version, tested against what?
- "We'll support CalDAV and Exchange" — no statement of which is first-class, who owns it, or what it costs per enterprise deal.
Appendices
Appendix A: Mechanics in Depth#
A.1 Range Expansion#
def instances(series, window_utc, tzdb):
z = tzdb.zone(series.tzid) # IANA rules for this zone
lo = z.to_local(window_utc.start) - series.duration # widen for overlap
hi = z.to_local(window_utc.end)
out = []
for local in rrule_iter(series.rrule, series.start_local, lo, hi):
if not valid_date(local): # Feb 30, Apr 31
continue # RFC default: omit
utc = z.resolve(local, gap="shift_forward", overlap="first")
orig = (series.id, local) # identity: original start
if local in series.exdates: continue
exc = exceptions.get(orig)
inst = apply(exc, series, utc) if exc else make(series, utc)
if overlaps(inst, window_utc): out.append(inst)
out += [make(series, z.resolve(d)) for d in series.rdates if in_window(d)]
out += exceptions.moved_into(series.id, window_utc) # by current start
return dedupe_by_original_start(out)
A.2 Free-Slot Search#
def free_slots(calendars, window, duration, working_hours):
busy = []
for cal in calendars: # parallel, one shard each
busy += busy_index.scan(cal, window) # [start_utc, end_utc)
busy += outside_working_hours(cal, window) # in each attendee's zone
busy.sort()
merged = merge_overlapping(busy) # O(n log n)
return [gap for gap in gaps(merged, window) if gap.length >= duration]
Working hours are wall-clock ranges in each attendee's zone, so they are converted per day — a 09:00–17:00 Berlin day and a 09:00–17:00 New York day overlap by a different number of hours in the March weeks when only one side has switched.
A.3 TZ Rebase#
Room bookings are re-derived inside the no-overlap constraint; a rebase can create a conflict when a booking anchored in a changed zone moves onto one anchored in an unchanged zone (a London-organized meeting in a room also booked from a Cairo-anchored series, say). Conflicts are reported to the organizers, never silently resolved.
Appendix B: Data Model#
calendars(calendar_id PK, owner_kind, owner_id, default_tzid, home_region)
events(event_id PK, calendar_id, ical_uid UNIQUE, start_local, duration,
tzid NULL, -- NULL = floating
all_day bool, start_date, end_date, -- dates for all-day, half-open
rrule NULL, rdates[], exdates[], until_utc,
sequence int, split_from NULL, title, location, visibility, status)
exceptions(event_id, original_start_local, overrides_json, cancelled,
current_start_utc, PK(event_id, original_start_local))
attendance(event_id, attendee_kind, attendee_id, partstat, role,
personal_reminders_json, hidden, last_seen_sequence,
PK(event_id, attendee_kind, attendee_id))
busy_intervals(calendar_id, start_utc, end_utc, event_id, original_start_local,
transparency, tzdb_version) -- projection, 18 months
room_bookings(room_id, during tstzrange, event_id, original_start_local)
-- EXCLUDE USING gist (room_id WITH =, during WITH &&)
reminder_timers(bucket_minute, user_hash_mod_64, user_id, event_id,
original_start_local, offset_min, channel, sequence)
change_log(calendar_id, seq, event_id, op, at) -- 30-day retention
The shard key is calendar_id for everything owned by a calendar; attendance rows live with the event (organizer's shard) and are projected into each attendee's change log and busy index on their own shards. See Schema Design for the access-pattern-first approach and Partitioning for co-location.
Appendix C: Coordination Mechanisms#
| Mechanism | Used For | Why Not Something Else |
|---|---|---|
| Single-shard transaction on organizer's calendar | Event + overlays + change log append | Writes stay local; attendees' views are projections |
Optimistic concurrency on SEQUENCE (If-Match) | Concurrent organizer edits (assistant + exec) | Conflicts are rare and must be shown to humans |
| Exclusion constraint on room ranges | Room and booking-page slots | The store enforces the invariant; no lock service |
| Idempotency keys on bookings and RSVPs | Retries from mobile and email gateways | See Idempotency |
| Change stream (Kafka, keyed by event_id) | Projections in order per event | Per-event ordering is enough; no global order needed |
| iTIP SEQUENCE ordering | External and cross-region copies | The standard already defines it |
Appendix D: Sync Protocol and Client Behavior#
- Tokens are per calendar, encode
(calendar_id, seq), and are valid for 30 days of change-log retention; older tokens get410with a pacedRetry-After. - Push is a hint with no payload; the device always pulls with its token. Channels expire and are renewed by the client.
- External CalDAV clients use the RFC 6578 sync-collection REPORT against the same change log, and receive whole iCalendar objects per changed resource; the adapter must round-trip RRULE, exceptions and VTIMEZONE without loss.
- Offline edits carry the SEQUENCE the device last saw; the server rejects stale organizer edits with a conflict the user resolves, and always accepts RSVP overlay writes (last writer wins on the attendee's own row).
- Local alarms: devices schedule 7 days of alarms from synced data and report the SEQUENCE scheduled per event so the server can suppress duplicate pushes.
Appendix E: Observability#
| Metric | Alert | Why |
|---|---|---|
cal.tzdb.release_ingest_lag_hours | > 1 h | Correctness dependency is behind |
cal.tzdb.rebase_lag_hours | > 6 h | Future instances still on old rules |
cal.tzdb.version_skew by zone and platform | Weekly report; page on changed zones | Devices disagree with the server |
cal.reminder.lateness_p99 by minute-of-hour | > 30 s | The :50 herd or a projection lag |
cal.freebusy.staleness_s p99 | > 5 s | People double-book what they saw free |
cal.freebusy.divergence (reconciler) | > 0.1% | Silent projection bugs |
cal.room.overlap_count | > 0 | A side-door write bypassed the constraint |
cal.sync.full_resyncs_per_min | > 3× baseline | Token invalidation storm |
cal.series.exceptions_dropped | Any spike | Destructive series edits |
cal.render.server_client_mismatch | > 0.01% | Recurrence or tz disagreement across clients |
Correctness vs availability: most calendar alerts here are correctness signals that would never show up as error rates. Page on rebase lag, divergence and room overlaps at all hours; they are the incidents users experience as "I was in the wrong place at the wrong time".
Appendix F: Scale Evolution#
| Stage | What Works | What You Add |
|---|---|---|
| < 100K users | Integrate with existing providers; no calendar of your own | Short-TTL availability cache; sync-token handling |
| 100K–5M | Own calendar on one PostgreSQL cluster; expand on read; timers in a queue | Intent-based time model from day one; room constraint |
| 5M–100M | Sharded by calendar; busy index; change-log sync; projections | tzdb automation; broadcast events; paced resync |
| 100M+ | Multi-region home placement; device-side expansion and alarms | Company-wide time standard; conformance suite; skew dashboards |
What you don't build on day one: your own tz database, a custom recurrence format, smart time suggestions, cross-tenant federation, or a materialized busy index for calendars nobody queries.