System design interview guide
Design a Hotel Reservation System: Inventory, Double Booking, Holds and Overbooking
Two courses by the author of this page:
770 lessons · 18 free to read
₹499 in India · $49 elsewhere, once
Get System DesignYou own this course
204 lessons · 10 free to read
₹999 in India · $49 elsewhere, once
Get AI EngineeringYou own this course
A hotel has a fixed number of rooms of each type for each night. A guest who books a King room for two nights takes one King for each of those nights. When the last King on a busy night is left, two guests can press Book within one second. Exactly one of them may get it, unless the hotel has decided to sell a few rooms more than it has. Everything hard about this design comes from that one small number per night: it must never go wrong, and many people want to change it at once.
A hotel reservation system keeps one inventory row per hotel, per room type, per night. Each row holds two counts: how many rooms of that type exist, and how many are already reserved. Booking a stay adds 1 to the reserved count of every night in the stay, inside one database transaction. A transaction is a group of changes that the database applies all together or not at all. To stop double booking, the check and the write must happen as one step. The simplest way is one guarded UPDATE: add 1 only if reserved + 1 is still within the limit, written as a WHERE condition in the SQL statement. If any night is full, the update changes fewer rows than the stay has nights, and the whole booking is rolled back (undone). A row lock (SELECT ... FOR UPDATE, which makes other bookings for those rows wait) or a version number check also works. The booking API is idempotent: the app sends a unique key with each booking, so a retry after a network error returns the first answer instead of holding a second room. A new booking starts as HELD with a timer of a few minutes. It becomes CONFIRMED when the card is authorized, or EXPIRED, which gives the room back. Hotels may allow overbooking, a set percentage above the real room count, based on how many guests usually do not arrive. Browsing reads a cache, a fast copy kept in memory that may be a few seconds old; only the inventory database decides a booking. Data is sharded by hotel_id, meaning it is split across database servers by hotel, so every row one booking touches lives on one server.
Where it shows up
This is a common system design question for backend roles, sometimes worded as "design a hotel booking system" or "design Booking.com". Its core shows up in any product that sells a limited thing for a date or a time: flight seats, event tickets, restaurant tables, doctor appointments, rental cars. If you can count limited stock per date and stop two buyers from taking the last unit, you can answer all of them.
Spec sheeta Hotel Reservation System
- 01Hotels and rooms
- 50,000 hotels, 5 million rooms
- 02Inventory rows
- about 91 million rows, about 9 GB
- 03Room nights and bookings per day
- 3.5 million room nights, 1.75 million bookings
- 04Booking writes per second
- about 20 average, 200 at peak
worked through below, with the maths
Why this question is asked
The question looks like a simple app that saves and reads records, and that is where candidates slip. A weak answer draws boxes for search, booking and payment, and stores a booked flag on each room. A strong answer finds the real problem in the first five minutes: a small number per night that many people try to change at once, and that must never be wrong. The interviewer then pushes on five points. How do you model inventory, per room or per room type? How do you stop two guests taking the last room, and what happens to the one who loses? What happens when the app retries a booking after a timeout? What happens to a room that a guest starts to book and then abandons? And why would a hotel ever sell more rooms than it has? Each one has a clear right answer and a few wrong ones that sound reasonable, so the question separates people who have built transactional systems from people who have only read about them.
This page is the free part.
The course goes deeper on the Hotel Reservation System design
₹499 in India$49 everywhere elseonce, for the whole course
The System Design course covers the Hotel Reservation System design across a run of lessons, not one page. These 4 alone are about 103 minutes of step-by-step reading, every one with a quiz.
- ACID Properties
Database transaction guarantees that keep your data correct.
- Serializability
Transactions execute as if they ran one at a time, the gold standard of isolation.
- Versioned Data
Tracking every change to a record, audit trails, undo history, and time-travel queries.
- Idempotency
Why some operations are safe to retry and others charge you twice, and how to make everything safe.
all part of
System Design Masterclass
770 lessons · about 283 hours · 18 free to read
₹499in India, by UPI
$49everywhere else, by PayPal
The concepts were explained in a clear and structured way, with practical examples that made even complex system design topics easier to understand. I especially liked the focus on real-world architecture, scalability, trade-offs. The content was well designed, engaging, and highly useful for anyone looking to strengthen their system design skills. Highly recommended for software engineers preparing for system design interviews or wanting to build a stronger foundation in designing scalable systems.
You own this course
Continue with ACID Propertiespay once,
yours for life
Live from the course bench
Before you design Hotel Reservation System, try the questions it rests on.
Real interview questions, live diagrams and measured lab results, pulled at random from the lessons. A new set every time you come back.
Requirements
Always clarify these in the first 5 minutes of the interview. Do not start drawing boxes until both lists are agreed.
Functional requirements
- Show which room types are free at a hotel for a date range, with the price for each night
- Let a guest reserve one or more rooms of a type for a date range, without ever selling more than the allowed limit
- Hold the rooms for a few minutes while the guest pays, and release them if payment does not finish in time
- Authorize the card at booking, and capture the money at check-out or as the hotel's policy says
- Let a guest cancel, following the hotel's cancellation policy, and give the rooms back
- Let hotel staff change room counts and prices, and set an overbooking limit per room type
- Make every booking request safe to retry: a request sent twice creates one reservation
Non-functional requirements
- Zero double bookings beyond the hotel's chosen overbooking limit, even with many guests booking the last room at once
- Strong consistency on the booking path: the inventory database is the only source of truth for what is left
- Browsing can show data a few seconds old, as long as the final booking step checks the database
- Availability pages answer fast (a few hundred milliseconds), because most traffic is people looking, not booking
- No lost money and no double charges: every payment call is safe to retry
- A slow or failed payment provider must not leave rooms blocked forever
- One hotel's busy night must not slow down bookings at every other hotel
Back-of-envelope scale estimates
Show your math. Pulling numbers from thin air signals you have not thought about the load.
Hotels and rooms
50,000 hotels, 5 million rooms
Assume 50,000 hotels on the platform with 100 rooms each on average: 50,000 x 100 = 5,000,000 rooms. Assume 5 room types per hotel (King, Twin, Suite and so on): 50,000 x 5 = 250,000 room types. These are interview assumptions. Say them out loud before you use them.
Inventory rows
about 91 million rows, about 9 GB
One row per room type per night, for a year ahead: 250,000 x 365 = 91,250,000 rows. At about 100 bytes a row that is about 9 GB. This fits on one database server, so size is not the reason to shard. Many writes aiming at a few popular rows, and the growing reservations table, are.
Room nights and bookings per day
3.5 million room nights, 1.75 million bookings
Assume 70% of rooms are sold each night: 5,000,000 x 0.70 = 3,500,000 room nights a day. Assume an average stay of 2 nights: 3,500,000 / 2 = 1,750,000 bookings a day. As a sanity check, Booking Holdings reported 1,144 million room nights booked in 2024, which is 1,144 million / 365 = about 3.1 million a day. So this is the size of the largest online travel group, a fair upper end for an interview.
Booking writes per second
about 20 average, 200 at peak
1,750,000 / 86,400 seconds = about 20 bookings a second. Assume peaks of 10 times the average: about 200 a second. Assume half of all holds are abandoned, so about 400 holds a second at peak. Each hold writes about 4 rows (2 nights of inventory, 1 reservation, 1 idempotency key): about 1,600 row writes a second. One well-sized database can do this. The hard part is that many of those writes aim at a few popular rows.
Availability reads per second
about 2,000 average, 20,000 at peak
Assume 100 availability checks for every completed booking (people compare many hotels and dates). 1,750,000 x 100 = 175 million checks a day. 175,000,000 / 86,400 = about 2,000 a second, and about 20,000 at a 10 times peak. This is why reads go to a cache and replicas, and only the booking step goes to the main database.
Reservations stored
about 640 million rows, about 320 GB a year
1,750,000 x 365 = 638,750,000 reservations a year. At about 500 bytes each (guest, dates, price, status, history) that is about 320 GB a year, and it keeps growing. This table, more than inventory, is what pushes the design toward sharding and archiving old stays.
Holds alive at once
about 240,000 at peak
400 new holds a second at peak, each living up to 10 minutes (600 seconds): 400 x 600 = 240,000 holds alive at once, in the worst case. Each one is a reservation row with status HELD and an expiry time, so the expiry job must be cheap: an index on (status, hold_expires_at).
High-level architecture
Draw two paths. The read path serves people who are looking. A guest searches a city and dates. A hotel search service (often a search engine such as Elasticsearch) finds candidate hotels. An availability service then answers, for each hotel and date range, which room types are free and at what price. It reads from a cache: a fast in-memory copy of rooms left per hotel, room type and night, kept in a store such as Redis. The cache is filled from the inventory database through change data capture (CDC), a feed of every committed change read from the database log, or simply expires after a few seconds. The page may say "2 rooms left" when one has just gone. That is fine, because nothing is promised until the write path runs. The write path serves people who are booking. The guest presses Book. The app sends POST /reservations with an Idempotency-Key header. The reservation service opens one transaction on the inventory database. It records the idempotency key, adds 1 to the reserved count of every night of the stay with a guarded UPDATE, and inserts a reservation row with status HELD and an expiry time 10 minutes ahead. Then it commits. Next, it asks a payment provider such as Stripe to authorize the card, which reserves the money on the card without taking it yet. When the authorization succeeds, a second small transaction moves the reservation from HELD to CONFIRMED. An expiry job finds HELD reservations whose time has passed, marks them EXPIRED and gives the nights back. Events such as "reservation confirmed" are written to an outbox table inside the booking transaction, then sent to a message queue for emails, the hotel's own systems and analytics. The inventory database is sharded by hotel_id, so the inventory rows, the reservation and the key for one booking always live on one server, and one ordinary transaction covers all of them.
In a real interview, sketch this on the whiteboard before diving into any single box.
One booking, from Book to confirmed
Follow one guest booking a King room at hotel 42 for the nights of Oct 11 and Oct 12. This is a good order to explain it in an interview.
- 1
Browse
The guest picks dates. The availability service reads the cache and shows King rooms as free, with the price per night. Nothing is reserved yet.
- 2
Send with a key
The guest presses Book. The app creates one random Idempotency-Key for this checkout and sends POST /reservations with hotel 42, room type King, check-in Oct 11, check-out Oct 13.
- 3
Route
The reservation service hashes hotel_id 42 to find its shard, the database server that holds hotel 42's rows.
- 4
Check the key
Inside one transaction, it inserts the key into the idempotency table. If the key is already there, this is a retry, and the service returns the stored reply without touching inventory.
- 5
Hold the nights
A guarded UPDATE adds 1 to reserved for King on Oct 11 and Oct 12, only where reserved + 1 stays within the limit. If it changes 2 rows, every night had room. If it changes fewer, the transaction rolls back and the guest is told the King is sold out for those dates.
- 6
Write the reservation
Still in the transaction, it inserts the reservation with status HELD, the price, and hold_expires_at 10 minutes from now. It saves the reply next to the key and commits.
- 7
Authorize the card
The service asks the payment provider to authorize the total, with the reservation id as the payment's own idempotency key. Authorize means the money is reserved on the card but not taken.
- 8
Confirm
On success, it updates the reservation to CONFIRMED, only WHERE status is still HELD. It writes a "confirmed" event to the outbox inside that transaction. Email and the hotel's systems receive it from the queue.
- 9
Or expire
If payment never finishes, the expiry job marks the reservation EXPIRED, only WHERE it is still HELD and its time has passed, and subtracts 1 from both nights. The King is bookable again.
- 10
Stay
At check-in the hotel assigns a room number. At check-out the money is captured, and the reservation becomes CHECKED_OUT.
Core components
Walk through each service. The interviewer wants to hear what each one owns, not just the names.
Availability service (read path)
Answers "which room types are free at this hotel for these dates, and at what price". It reads rooms left per night from the cache and the price from the rate tables. A room type is free for a stay only if every night of the stay has room, so it takes the smallest count across the nights. It never decides a booking; it only shows a recent picture.
Reservation service (write path)
The only component allowed to change inventory. It runs the booking transaction: record the idempotency key, guarded UPDATE on each night, insert the reservation. It also runs confirm, cancel and check-in, each as one guarded state change. It holds no state of its own, so you run several copies behind a load balancer, a server that spreads requests across them.
Inventory database
A relational database such as PostgreSQL or MySQL, because the booking needs transactions, row locks and CHECK constraints. It holds the inventory rows, reservations, idempotency keys and the outbox, all sharded by hotel_id. A shard is one database server that holds part of the data. Each shard has a primary that takes writes and replicas that copy it for reads.
Availability cache
An in-memory store (for example Redis) with keys like hotel:42:king:2026-10-11 and the number of rooms left. It is updated from the database's change feed and expires after a short time, so a missed update heals by itself. Because it can be wrong for a few seconds, the write path never reads it.
Expiry job
Runs every few seconds. It finds reservations with status HELD and hold_expires_at in the past, moves each one to EXPIRED with a guarded update, and in that one transaction gives its nights back. Several copies can run at once, because the guarded update lets only one of them win for each reservation.
Payment integration
Talks to a payment provider. At booking it authorizes the card (reserves the money); at check-out, or by the hotel's policy, it captures it (takes the money). Every call carries an idempotency key, so a retry after a timeout cannot charge twice. Payment results come back by webhook, a call the provider makes to our server, and are applied with a guarded state change too.
Outbox and message queue
The outbox is a table where the booking transaction also writes an event, such as "reservation 555 confirmed". A relay reads new outbox rows and publishes them to a queue such as Kafka. Because the event is written inside the booking transaction, there is never a booking without its event, or an event without its booking. Emails, the hotel's property management system (the software the front desk uses) and analytics read from the queue.
Hotel admin and rate service
Where hotel staff set how many rooms of each type exist, the price for each night, the cancellation policy and the overbooking percentage. Lowering the room count below what is already reserved must be refused or flagged, because the CHECK constraint would reject it.
Data model
Pick the right store per table. Justify each choice with the access pattern, not by reflex.
room_type_inventoryhotel_idroom_type_idstay_datetotal_inventorytotal_reservedoverbook_pctThe heart of the design. Primary key (hotel_id, room_type_id, stay_date). total_reserved counts held and confirmed rooms. A CHECK constraint keeps total_reserved within total_inventory plus the overbooking percentage, so even a bug in the app cannot oversell. A stay of 3 nights touches 3 rows. stay_date is the hotel's local date, so a guest in another time zone still books the right night.
reservationsreservation_idhotel_idroom_type_idguest_idcheck_incheck_outroomsstatushold_expires_attotal_priceOne row per booking. status is one of HELD, CONFIRMED, EXPIRED, CANCELLED, NO_SHOW, CHECKED_IN, CHECKED_OUT. check_out is the morning the guest leaves, so the nights are check_in up to the day before check_out. Index (status, hold_expires_at) for the expiry job. Stored on the hotel's shard.
idempotency_keysguest_ididem_keyrequest_hashreservation_idresponse_coderesponse_bodycreated_atPrimary key (guest_id, idem_key). Written inside the hold's transaction. request_hash detects a reused key sent with a different body, which is an error. Old rows can be deleted; Stripe allows removing keys once they are at least 24 hours old.
rateshotel_idroom_type_idstay_datepricecurrencyThe price per room type per night. Kept apart from inventory because prices change often and are read far more than they are written. The total price is computed at booking and stored on the reservation, so a later price change never alters a booking already made.
paymentspayment_idreservation_idprovider_refkind (authorize / capture / refund)amountstatusOne row per call to the payment provider, with the provider's own id. Lets support answer "was this guest charged?" and lets a job find authorizations that need to be captured or cancelled.
outboxevent_idhotel_idtypepayloadcreated_atpublished_atEvents written in the booking transaction and sent to the queue by a relay. published_at is set once the queue has accepted the event.
Deep dives
These are the conversations the interviewer is steering you toward. Practice each one until you can talk through it without notes.
Inventory: one row per hotel, room type and night
The first design choice decides everything after it. Guests do not book room 512. They book "a King room for two nights", and the hotel picks the room number later. So the system should count rooms of each type for each night, instead of tracking each physical room. The table is room_type_inventory, with one row per hotel, room type and date. Each row holds total_inventory, the number of rooms of that type that exist, and total_reserved, the number already held or confirmed. A stay from Oct 11 to Oct 13 covers the nights of Oct 11 and Oct 12, so it touches two rows. The check-out day is not a night of the stay. Counting per type has three benefits. Each night is one number to check, not a hundred rows to search. The number fits a database constraint, a rule the database itself enforces on every write. And hotels already think and sell this way. A CHECK constraint is that rule here: total_reserved may never pass the limit. Even if a bug in the application skips a check, the database refuses the write. Per-room rows make sense in a different product, such as a meeting-room booking tool where the user picks a specific room. For that case PostgreSQL offers an exclusion constraint on a date range, which refuses two overlapping bookings of one room.
CREATE TABLE room_type_inventory (
hotel_id BIGINT NOT NULL,
room_type_id BIGINT NOT NULL,
stay_date DATE NOT NULL,
total_inventory INT NOT NULL, -- rooms of this type
total_reserved INT NOT NULL DEFAULT 0, -- held + confirmed
overbook_pct INT NOT NULL DEFAULT 0, -- 5 = sell up to 105%
PRIMARY KEY (hotel_id, room_type_id, stay_date),
CHECK (total_reserved * 100 <= total_inventory * (100 + overbook_pct))
);
UPDATE room_type_inventory
SET total_reserved = total_reserved + 1
WHERE hotel_id = 42 AND room_type_id = 7
AND stay_date >= '2026-10-11' AND stay_date < '2026-10-13'
AND (total_reserved + 1) * 100 <= total_inventory * (100 + overbook_pct);
-- 2 rows changed: book it. Fewer than 2: ROLLBACK, sold out.
The double-booking race
Here is the bug most first designs have. The code reads the count, checks it in the application, then writes the new count. Two guests want the last King on Oct 11, with 99 of 100 taken. Guest A reads 99. Guest B reads 99. A sees room and writes 100. B also saw room and writes 100. Two guests now hold one room, and the stored count of 100 even looks correct, so nothing alerts anyone. This is called a race condition: the result depends on the timing of two requests that run at once. Adding a cache lock or a check in the app does not fix it, because the check and the write are still two separate steps that another request can get between. The fix is to make the check and the write one step that the database itself protects. PostgreSQL's default isolation level (the rules for how transactions running at once see each other's changes) is called Read Committed, and in it a guarded UPDATE does this. When B's UPDATE finds a row that A has already changed but not committed (made final), B waits. When A commits, PostgreSQL re-checks B's WHERE clause on the new version of the row. 100 + 1 is more than 100, so B's update changes 0 rows, and the service tells B that the room is gone. A booking for several nights must change one row per night. If the count of changed rows is short, one night was full, and the whole transaction is rolled back so no partial hold is left.

Three ways to stop it: lock, version, or guarded write
Interviewers usually ask you to compare the options, so know all three. The first is a pessimistic lock: assume a clash will happen and lock first. SELECT ... FOR UPDATE reads the inventory rows and locks them. Any other transaction that tries to update or lock those rows waits until this one commits or rolls back. The application checks the counts at leisure, then updates. It is simple to reason about, and it suits a hotel where many guests fight for one night. Lock the nights in date order every time. If one booking locked Oct 12 then Oct 11 while another locked Oct 11 then Oct 12, each could wait for the other forever, which is called a deadlock. The second is an optimistic version check: assume clashes are rare and detect them at write time. Each row has a version number. The app reads the row and its version, then writes WHERE version is still the one it read. If someone changed the row in between, 0 rows change, and the app reads again and retries. No locks are held while the guest's request is being processed, which is good when collisions are rare, and wasteful when they are common, because every loser retries. DynamoDB's conditional writes work this way: the write succeeds only if a condition on the item is still true. The third is the guarded write from the previous section: put the business rule itself in the WHERE clause. The loser gets 0 rows and a clear answer, sold out, with no retry needed. Back all three with the CHECK constraint, so a bug cannot oversell. A fourth option, raising the isolation level to Serializable, also works in PostgreSQL, but the database then aborts one of two clashing transactions with a serialization error, and the application must retry the whole transaction. The course's Serializability lesson walks through pessimistic locking and optimistic validation side by side, with this retry rule.
-- 1. Pessimistic: lock the nights, then check and write
BEGIN;
SELECT stay_date, total_inventory, total_reserved, overbook_pct
FROM room_type_inventory
WHERE hotel_id = 42 AND room_type_id = 7
AND stay_date >= '2026-10-11' AND stay_date < '2026-10-13'
ORDER BY stay_date -- same order every time: no deadlock
FOR UPDATE; -- other bookings for these rows wait
-- the app checks every night has room, then UPDATE ... and INSERT ...
COMMIT; -- locks released
-- 2. Optimistic: write only if nobody changed the row since we read it
UPDATE room_type_inventory
SET total_reserved = 88, version = version + 1
WHERE hotel_id = 42 AND room_type_id = 7 AND stay_date = '2026-10-11'
AND version = 13; -- 0 rows: read again and retry
An idempotent booking API
Networks fail in an annoying way: the app cannot tell whether its request was lost or only the reply was lost. If the reply was lost, the booking already happened. A retry without protection would hold a second room and, later, charge the card twice. An operation is idempotent when doing it twice has the effect of doing it once. To make booking idempotent, the app generates a random key, such as a UUID, when the guest starts checkout, and sends it in an Idempotency-Key header with the booking and with every retry of it. The server stores the key, a hash of the request body and the reply. Stripe's API works this way: it saves the status code and body of the first request for each key and returns them for later requests with that key, it compares the parameters and returns an error if they differ, and keys can be removed once they are at least 24 hours old. The key must be written inside the database transaction that makes the hold. Then a hold without its key, or a key without its hold, cannot exist. If two copies of the request arrive at once, the primary key lets only one of them insert the key. The other is treated as a retry; if the first has not finished yet, it gets a "try again shortly" reply. Use the reservation id as the idempotency key for the call to the payment provider too, so a retried authorization does not reserve the money twice. The course's Idempotency Keys lesson replays this lost-reply case step by step.
CREATE TABLE idempotency_keys (
guest_id BIGINT NOT NULL,
idem_key TEXT NOT NULL, -- made by the app, once per checkout
request_hash TEXT NOT NULL, -- same key, other body: reject
reservation_id BIGINT,
response_code INT,
response_body JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (guest_id, idem_key)
);
BEGIN;
INSERT INTO idempotency_keys (guest_id, idem_key, request_hash)
VALUES (9001, 'c1f7e2...', 'sha256:ab12...')
ON CONFLICT (guest_id, idem_key) DO NOTHING;
-- 0 rows inserted: a retry. ROLLBACK and return the stored reply.
-- 1 row: guarded UPDATE on inventory, INSERT reservation,
-- UPDATE idempotency_keys SET reservation_id, response_...
COMMIT;
The reservation state machine and hold expiry
A guest who presses Book still has to pay, and that can take minutes. If the room is only counted after payment, two guests can both reach the payment page for the last room, and one pays for nothing. So the booking first creates a hold: the nights are counted as reserved and the reservation gets status HELD with hold_expires_at set, here 10 minutes ahead. A state machine lists every state a reservation can be in and the only moves allowed between them. HELD becomes CONFIRMED when the card is authorized in time, EXPIRED when the timer runs out, or CANCELLED when the card is declined. CONFIRMED becomes CANCELLED, NO_SHOW or CHECKED_IN. The rule that makes this safe is simple: every move is an UPDATE that names the state it expects to leave, such as WHERE status = 'HELD'. Now consider the nasty case. The payment succeeds at 9 minutes 59 seconds, but the success message reaches the service at 10 minutes 2 seconds, after the expiry job ran. The expiry job's update matched first, so the confirm update finds no HELD row and changes 0 rows. The service then cancels the card authorization so the guest is not charged, and tells the guest to try again. Without the guarded update, both jobs would win: the room would be given back and also confirmed. Expired holds give their nights back inside the transaction that marks them EXPIRED, so the count and the status can never disagree. Several expiry workers can run at once, because the guarded update lets only one of them win for each reservation.
-- Expiry job, for one reservation
BEGIN;
UPDATE reservations SET status = 'EXPIRED'
WHERE reservation_id = 555 AND status = 'HELD'
AND hold_expires_at < now();
-- 0 rows: it was confirmed in time. COMMIT and stop.
-- 1 row: give the nights back
UPDATE room_type_inventory
SET total_reserved = total_reserved - 1
WHERE hotel_id = 42 AND room_type_id = 7
AND stay_date >= '2026-10-11' AND stay_date < '2026-10-13';
COMMIT;
-- Payment succeeded
UPDATE reservations SET status = 'CONFIRMED'
WHERE reservation_id = 555 AND status = 'HELD';
-- 0 rows: the hold already expired. Cancel the card authorization.
Payment: authorize at booking, capture later
A card payment can happen in two steps. Authorize reserves the amount on the guest's card; capture actually takes it. Stripe's documentation uses hotels as its example: a hotel often authorizes the full amount before the guest arrives and captures it at check-out. An authorization does not last forever. For most card networks, an online authorization is valid for 7 days. Some networks allow an extended authorization of up to 30 days for hotels and lodging. So for a stay booked three months ahead, the system cannot simply authorize at booking and capture at check-out. It has two common choices. For a prepaid rate, capture at booking and refund under the cancellation policy. For pay later, save the card with the provider at booking, check it, and authorize closer to arrival. The reservation service treats the provider as slow and unreliable. Every call carries an idempotency key. Results also arrive by webhook, and the webhook handler applies them with a guarded state change, so a result that arrives twice, or late, does no harm. A job looks for authorizations that are about to expire, and for reservations that are CONFIRMED with no capture after check-out. Money moves are recorded in the payments table, so support can answer "was I charged?" from our own data. The course's Design a Payment System capstone covers the ledger and reconciliation side in depth.
Overbooking on purpose
Hotels know that some booked guests will not come. A no-show is a guest with a confirmed booking who never arrives. Others cancel too late for the room to be sold again. An empty room that night earns nothing, so many hotels deliberately sell a few more rooms than they have. SiteMinder, which sells software to hotels, says hotels may overbook by a percentage "such as 2%, 3%, or even 10%", and that many set it at or just below their historical no-show and cancellation rate. So 10% is an upper example, and the real number differs by hotel and by season. The cost of getting it wrong is real. When more guests arrive than there are rooms, the hotel "walks" a guest: it arranges and pays for a room at another hotel, plus transport and often some compensation. In the system, overbooking is just one more number per room type, overbook_pct, set by the hotel's revenue team. The CHECK constraint and the guarded UPDATE use total_inventory x (100 + overbook_pct) / 100 as the limit. For 100 King rooms at 5%, that is 105. Use whole-number math, as in the code, so 100 x 1.05 never becomes 104.99999 in floating point. The system does not choose the percentage. It stores it, enforces it exactly, and reports how often guests were walked, so the hotel can tune it.
rooms, no_show = 100, 0.08
for pct in (0, 5, 10):
sold = rooms * (100 + pct) // 100 # the limit the CHECK enforces
arrive = sold * (1 - no_show) # expected, on an average night
print(pct, sold, round(arrive, 1), round(rooms - arrive, 1))
# 0 100 92.0 8.0 8 rooms empty
# 5 105 96.6 3.4 about 3 empty
# 10 110 101.2 -1.2 about 1 guest walked
Read path: fast and a little stale. Write path: always right
About 100 people look for every one who books, so the read path must be cheap. The availability service reads rooms left from a cache, keyed by hotel, room type and night. A cache is a fast copy of data kept in memory. The cache is updated from the database's change feed, and every entry also expires after a short time, so if an update is ever missed the entry heals itself. Hotel pages and photos are served from read replicas, copies of the database that take reads off the primary. The cost is staleness. A guest may see "2 rooms left" when one was booked a second ago. That is acceptable for one reason: the booking step never trusts the cache. It runs the guarded UPDATE on the primary, and if the room is gone, the guest sees a clear message and the cache entry is refreshed. The opposite design, reading the primary for every search, would put 20,000 reads a second on the servers that must also lock rows for bookings, and slow down the one path that must not be slow. One more source of change: hotels also sell their rooms on other websites. A channel manager is software that sends a hotel's rooms and prices to every website it sells on, and updates them all when one of them makes a sale. If those updates are slow, two websites can sell the last room, so the integration should push changes as they happen.
Sharding by hotel and the hot row
Sharding means splitting the data across several database servers, each called a shard. The shard key decides where each row goes. Here the key is hotel_id. A hash function turns the id into a number, and that number picks the shard. Inventory rows, reservations, idempotency keys and outbox events for one hotel all go to one shard. So a booking, which only ever touches one hotel, is one ordinary transaction on one server, and never needs a distributed transaction across servers. If you sharded reservations by guest instead, a booking would touch two shards, and you would need two-phase commit (a protocol where a coordinator asks every server to prepare, then to commit) or a saga (a chain of steps, each with an undo step). Both are slower and harder to get right. A guest's "my trips" page then needs a lookup by guest, served from a separate index table or a search copy, kept up to date from the outbox. Sharding does not fix the hot row. Hotel 42's King on New Year's Eve is one row, and every booking for it must wait for that row's lock. If each booking holds the lock for about 5 milliseconds, that row takes at most 1000 / 5 = 200 bookings a second, however many shards there are. That is far more than any real hotel sells in a second. For a flash sale of thousands of discounted rooms at once, put a waiting room or a queue in front, so requests arrive at the row at a pace it can handle. Keep the booking transaction short: no call to the payment provider while the lock is held.


Trade-offs to discuss
Every senior interviewer expects you to surface at least 3 of these. Pick the decisions, state the alternatives, and justify your choice.
Count per room type vs track each physical room
Counting per type gives one row and one number per night, which is easy to guard with a constraint and matches how hotels sell. Tracking each room lets a guest choose a specific room, but every booking must search for a free one and lock it. Use counts for hotels; use per-room rows with a date-range exclusion constraint for products where the user picks the exact room.
Pessimistic lock vs optimistic version vs guarded update
A lock is easy to reason about and fair under heavy contention, but every booking for that night waits in line. A version check holds no locks, but losers must retry, which wastes work when clashes are common. A guarded update needs no retry loop and gives the loser a clear answer, but only works when the rule fits in a WHERE clause. For hotel inventory it does, so it is the default here.
Hold the room before payment vs only after payment
Holding first means a guest who reaches the payment page will get the room, at the cost of rooms being blocked for a few minutes by people who leave, plus an expiry job. Counting only after payment is simpler, but two guests can pay for one last room and one must then be refunded. Hold first, with a short timer.
Read availability from a cache vs from the primary database
The cache is fast and keeps the primary free for bookings, but it can be a few seconds behind. Reading the primary is always exact, but it cannot serve about 100 lookers for every booker. Use the cache for browsing and the primary for the booking step only.
Shard by hotel_id vs by guest_id
By hotel keeps every row of one booking on one server, so one plain transaction is enough. By guest makes "my trips" a single-shard read, but every booking then crosses shards and needs a saga or two-phase commit. Shard by hotel and build the per-guest view as a separate index.
Allow overbooking vs never sell more than exists
Overbooking fills rooms that no-shows would leave empty, but sometimes a guest must be sent to another hotel, at the hotel's cost. Never overbooking is simpler and never walks a guest, but leaves money on the table every night. Make it a per-hotel setting, enforced exactly, and default it to 0.
Free PDF · 18 pages
Get the free System Design Interview Cheat Sheet
The interview in seven stages, the numbers worth knowing by heart, and twelve classic systems on one page each, every line linked to the lesson it comes from.
Follow-up questions to expect
After the main design, the interviewer usually picks one part and asks more. Here is what to say.
A guest books 3 rooms for 4 nights. What changes?
Add 3 instead of 1 in the guarded UPDATE, check reserved + 3 against the limit, and expect exactly 4 rows changed. If fewer than 4 change, roll back. The reservation row stores rooms = 3, and expiry or cancellation subtracts 3.
What if the expiry job is down for an hour?
Abandoned holds keep their rooms blocked, so the hotel looks fuller than it is, but nothing is double booked. Make reads treat a HELD row past its expiry as free, and have the booking path expire stale holds for that hotel before it gives up. Alert when the oldest HELD reservation is far past its expiry time.
Hotel staff lower the King count from 100 to 90, but 95 are already reserved. What happens?
The CHECK constraint rejects the update, which is what you want. The admin tool should show how many reservations are affected and ask staff to move or cancel some first, or raise overbook_pct for those nights on purpose.
How do you show a price that stays fixed during checkout?
Compute the total at the moment of the hold, store it on the reservation, and charge that amount. A price change after that point does not alter the reservation. If the price changed between browsing and booking, show the new total before asking for payment.
Why not use a Redis lock instead of database locks?
A Redis lock lives in a different system from the count it protects. If the lock expires while a slow request is still working, a second request can enter, and the database never checked anything. The guarded UPDATE and the CHECK constraint protect the count where it is stored. Redis is fine as a cache for reads.
The payment provider times out. Is the guest charged or not?
You do not know yet. Leave the reservation HELD, retry the call with its original idempotency key, and also wait for the provider's webhook. Whichever answer arrives first moves the state with the guarded update; a second answer changes 0 rows. If the hold expires first, cancel the authorization when it shows up.
How do you handle a guest who books one hotel on two websites?
They are two reservations, possibly through a channel manager, and both count against inventory. The system cannot know the guest meant one. The cancellation policy and no-show history handle it, and both feed the hotel's choice of overbooking percentage.
How a Hotel Reservation System actually does it
Every mechanism on this page is documented by the system that provides it. PostgreSQL's documentation says FOR UPDATE locks the selected rows so that other transactions trying to update, delete or lock them wait until the current transaction ends. It also explains that in Read Committed mode, a second updater waits for the first, then re-evaluates its WHERE clause on the updated row, which is why the guarded update is safe. The PostgreSQL range-types page even shows a room_reservation table with an exclusion constraint that refuses two overlapping bookings of one room. Stripe's idempotent requests guide describes saving the first response for a key, comparing parameters on retries, and removing keys once they are at least 24 hours old. Brandur Leach published a full Postgres design for idempotency keys, with a unique constraint on (user_id, idempotency_key) and a stored response. Stripe's guide to placing a hold uses hotels as its example, and its extended authorization page lists 30-day windows for hotel and lodging merchants on some card networks. Amazon's DynamoDB guide shows the optimistic version: a conditional write succeeds only if the item still matches an expected value. On the business side, SiteMinder describes overbooking of "2%, 3%, or even 10%", set at or just below a hotel's own no-show and cancellation rate, and explains walking a guest to another hotel. Its channel manager page explains how one hotel's rooms are kept in step across many booking websites. For scale, Booking Holdings reported 1,144 million room nights booked in 2024, about 3.1 million a day, which is why the estimates above use a few million room nights a day as the upper end.
Sources
- PostgreSQL documentation: explicit locking (FOR UPDATE)
- PostgreSQL documentation: transaction isolation (Read Committed re-check, Serializable retry)
- PostgreSQL documentation: constraints (CHECK, exclusion)
- PostgreSQL documentation: range types (room reservation exclusion constraint)
- Stripe API reference: idempotent requests
- Stripe documentation: place a hold on a payment method
- Stripe documentation: extended authorizations
- Brandur Leach: Implementing Stripe-like idempotency keys in Postgres
- Amazon DynamoDB developer guide: conditional writes
- SiteMinder: hotel overbooking, pros, cons and strategy
- SiteMinder: what a channel manager does
- Booking Holdings: fourth quarter and full year 2024 results (SEC filing)
Lessons to study before this interview
If any of these topics are fuzzy, the interviewer will catch it. Each lesson is 15 to 60 minutes with diagrams, code, and a quiz.
ACID Properties
foundation / core fundamentals
Serializability
advanced / consistency models
Versioned Data
foundation / database fundamentals
Idempotency
foundation / core fundamentals
Idempotency Keys
intermediate / api design protocols
Finite State Machine
advanced / distributed systems core
Distributed Locks
advanced / distributed systems core
Transactional Outbox Pattern
intermediate / messaging event systems
Change Data Capture (CDC)
intermediate / data replication distribution
Cache-Aside Pattern
foundation / caching strategies
Read Replicas
foundation / database fundamentals
Database Sharding
foundation / database fundamentals
Saga Pattern
advanced / distributed systems core
Design a Payment System
capstone / capstone
Frequently asked questions
Practice with 770 system design lessons
Lifetime access for ₹499 in India or $49 elsewhere. Interactive diagrams, quizzes, and 20 capstone projects to practise on.
