Market data14 min read

Case study: Airbnb ADR vendors and their systematic bias

This article discusses why scraped ADR figures should be treated with caution and where to be most skeptical. Not all ADR results are equally wrong: some are more reliable than others. The aim is to show which listings and markets produce the most distorted figures.

The cause is a fundamental limit in how the data is displayed on Airbnb's side, from orphaned nights to large gaps forming in calendars creating uncertainty.

How Airbnb occupancy and ADR data is collected

Airbnb doesn't publish bookings. Every third-party occupancy and revenue number, from every vendor alike, is inferred from the one thing that is public: the calendar. Each day a night is either open or closed. When an open night closes, the tool has to decide what happened.

There are three possible answers. A guest booked it. The host blocked it (personal use, maintenance, a stay taken on another channel). Or the night became impossible to book and was closed. The first earns money. The other two don't. The calendar looks the same in all three cases.

When a closed gap really is a booking

Take a listing with a 2-night minimum. Two bookings sit on the calendar with exactly two free nights between them. Tomorrow, both nights are closed.

1 Before: 2 open nights between two stays 9 10 11 12 13 14 open 2 A guest books both nights 9 10 11 12 13 14 new booking 3 What a data vendor counts 9 10 11 12 13 14 2 counted, 2 booked
Figure 1. Minimum stay 2 nights, gap of 2. The only booking that fits fills the whole gap, so the vendor count is right.

This is the case that makes calendar inference look reliable. The gap and the minimum stay match, so there is only one way the gap could have been booked. Now change the gap by one night.

Orphan nights: the gap a minimum stay can't fill

Same listing, same 2-night minimum, but the gap is three nights. A guest books two of them. The third is left alone between two stays. It's shorter than the minimum, so nobody can book it. The host, their channel manager or a gap-night rule closes it. It earns nothing.

1 Before: 3 open nights between two stays 9 10 11 12 13 14 15 3 open nights 2 A guest books 2 of them 9 10 11 12 13 14 15 new booking 1 night left: under the minimum 3 What a data vendor counts 9 10 11 12 13 14 15 booked counted as booked
Figure 2. Minimum stay 2 nights, gap of 3. Two nights booked, three counted. The sandwiched night earned nothing.

From the outside, the calendar went from three open nights to three closed nights. That's the same trace a 3-night booking leaves. The stranded night gets counted as occupancy and priced at the listing's rate. This kind of stranded night is usually called an orphan night, and hosts have been complaining about them for as long as minimum stays have existed. What gets less attention is what they do to market data.

How the minimum night stay scales the error

Raise the minimum to three nights and widen the gap to five. A guest books the three nights in the middle. One night is stranded on each side. Both close. Both look booked.

1 Before: 5 open nights between two stays 9 10 11 12 13 14 15 16 17 18 5 open nights 2 A guest books the 3 in the middle 9 10 11 12 13 14 15 16 17 18 new booking sandwiched sandwiched 3 What a data vendor counts 9 10 11 12 13 14 15 16 17 18 booked counted counted
Figure 3. Minimum stay 3 nights, gap of 5. Three nights booked, five counted: 40% of the gap is error.

The pattern generalises. When a single stay lands in a gap that can't hold two stays, every night it doesn't use is dead:

D=G−L Dmax=2(m−1)

D: nights stranded by one stay; G: gap length; L: stay length; m: minimum stay. The ceiling is 2 nights per stay at a 2-night minimum, 4 at 3 nights and 58 at 30 nights.

Two phantom nights per stay at a 2-night minimum sounds small, and on one listing it is. But the ceiling rises in step with the minimum stay, and there's one market where the minimum stay is set by law.

New York's 30-night minimum

Since New York City began enforcing Local Law 18 in September 2023, stays under 30 days need a registered host who is present during the stay. Most whole-home listings in the city moved to a 30-night minimum as a result. Run the same arithmetic there.

existing booking30-night bookingsandwiched: can't be bookedcounted as bookedcounted, nobody stayed

What happened

November 2026 S M T W T F S 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30
December 2026 S M T W T F S 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31

One 30-night booking, Nov 17 – Dec 16. The 14 nights before it and 15 after it are each under 30: sandwiched

What a data vendor counts

November 2026 S M T W T F S 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30
December 2026 S M T W T F S 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31

59 nights counted as booked. 29 of them (ringed) nobody stayed in

Figure 4. New York, 30-night minimum, gap of 59. One 30-night booking in the middle of the gap. Every night left on either side is under 30, so none of them can be booked: 30 nights booked, 59 counted.

That's 29 nights counted as booked that nobody booked, from a single gap on a single listing, and it isn't even the worst case: with 29 stranded nights possible on each side, one booking can hide up to 58. Real calendars won't hit the ceiling every time. But the direction of the error never changes. Orphan nights can only turn into phantom bookings. There is no mechanism that turns a real booking into an orphan night. So the error doesn't average out across a market. It stacks.

A 30-night minimum also creates a second trap. Any gap shorter than 30 nights is dead from the moment it appears. If the host closes it, and many do to keep the calendar tidy, it enters the data as a booking of up to 29 nights that never happened.

When a whole-month rule removes the problem

Some New York hosts go one step further than the 30-night minimum: guests can only book whole calendar months, starting on the 1st. Then every booking is one or more complete months, and anything left open is one or more complete months too. No gap shorter than the minimum can ever form, so nothing gets sandwiched. For those listings the calendar is an honest record, and occupancy, ADR and revenue figures built from it are clean.

The one month to watch is February. At 28 nights it falls under a 30-night minimum, so a host has to sell it as part of a longer stay or set a separate minimum for it. If neither happens and January and March are booked, February sits empty and unbookable between them.

Why the busiest Airbnb listings are overestimated the most

An orphan night isn't a one-off accident. It's a risk that comes with every booking. Each new stay lands somewhere in a gap, and each landing is another chance to strand the nights around it. A listing that takes 10 stays a year runs that risk 10 times. A listing that takes 50 runs it 50 times. The more often the event repeats, the more often the error happens, and the errors only ever add up in one direction.

Dyear≈S×p×d‾

S: stays per year; p: chance a stay strands nights; d: nights stranded when it does.

If that middle term were fixed, the error would grow in step with bookings. It isn't fixed. It rises with demand too. A quiet listing has long open stretches, and a stay dropped into a 40-night gap leaves plenty of room on both sides. A busy listing's calendar is chopped into short gaps between stays, which is exactly where orphans form. So busier listings have more events and a higher strike rate on each one.

We ran a simple simulation to see how fast this compounds: one year of calendar, booking requests at random dates and lengths, and a request accepted only if its nights are free. Any leftover run shorter than the minimum stay closes.

The quiet listing's calendar is almost clean: about 1% of the nights it appears to have booked were orphans. The busy listing is overstated by more than one night in ten. Bookings went up about 7×, and the phantom nights went up more than 60×. That's the compounding: the error grows faster than the bookings that cause it.

What this does to ADR and revenue

Occupancy takes the hit directly: every phantom night is counted as a booked night. Revenue follows, because each phantom night is valued at the price the calendar was showing for it. ADR can drift up as well. A stranded night never goes through the last-minute discount that often gets a real night booked. It stays at its full asking price until it closes, and then it's counted at that price. The busier the listing, the more of these full-price phantom nights get mixed into its average.

Now put that next to how market benchmarks get built. Demand on Airbnb converges on well-reviewed listings. They have the most calendar activity, which makes them look like the most informative data points. If an estimator leans on high-activity listings to calibrate occupancy and ADR, which is a natural thing to do since they carry the most signal, it's leaning on exactly the listings with the largest and fastest-growing error. The error then doesn't stay in those listings. It spreads into the benchmark every other listing is measured against. That's our reading of how the numbers behave, not something any vendor has published. But it fits the common complaint that projections run hottest for the listings that look most attractive.

It's also why the usual explanation for the gap falls short. Cleaning fees get blamed for inflated revenue figures, but fees can be read off the listing. Orphan nights can't.

Booked or blocked? One calendar, many histories

Go back to a 2-night minimum and a 5-night gap that closes overnight. Figure 5 lists six of the histories that could have produced that change:

existing bookingfirst bookingsecond bookingsandwichedcounted as booked
M T W T F S S M T 9 10 11 12 13 14 15 16 17 5 booked: one 5-night stay 9 10 11 12 13 14 15 16 17 5 booked: 3 + 2 9 10 11 12 13 14 15 16 17 5 booked: 2 + 3 9 10 11 12 13 14 15 16 17 4 booked, 1 orphan 9 10 11 12 13 14 15 16 17 4 booked, 1 orphan 9 10 11 12 13 14 15 16 17 3 booked, 2 orphans 9 10 11 12 13 14 15 16 17 Vendor: 5 counted
Figure 5. Minimum 2, gap of 5. Six histories, three different revenue outcomes, one identical calendar. Which one happened is a latent variable: it exists, but the calendar never reveals it.

Scraping more often helps less than you'd think. Daily snapshots can sometimes show the stays arriving one at a time, which rules some histories out. But when the last nights close, you still can't tell a short booking from an orphan being shut. A host might also turn down a 3-night request that would strand two nights, or accept a 1-night exception below their own minimum. Both happen. Neither leaves a trace.

This is what irreducible error means. It isn't noise that more data smooths out, or a bias a smarter model can learn to correct listing by listing. The fact that would settle the question, whether a guest paid for that night, is not in the public data at all.

How Airbnb data vendors handle booked vs blocked nights

Vendors that sell Airbnb occupancy, ADR and revenue data from scraped listings describe booked versus blocked as a classification problem: gather signals such as length of stay, lead time and reviews, train a model, and label each closed night. Some also take bookings directly from property managers for part of the market, and for those listings they can see real reservations. For the rest, the published methodology pages we reviewed don't explain how nights sandwiched by a minimum stay are treated.

Classification assumes the answer is somewhere in the data, waiting to be found. For sandwiched nights, it isn't. A booked night and a night closed by the minimum-stay rule leave the same trace on the public calendar.

A better estimate: treat the calendar as constraints

If the question can't be answered night by night, stop asking it night by night. The honest approach is to treat the calendar as a set of constraints and reason over what they allow.

The minimum stay is the most useful constraint there is, because it's a hard rule. For a gap of known length and a known minimum, you can list every combination of stays and orphan runs that closes it. You can attach a prior to each one: a single long stay is more likely than three short ones back to back, and hosts rarely break their own minimum. The expected number of booked nights is then a weighted average:

E[B]=∑h∈HP(h)·B(h)

H: every way the gap could have closed under the minimum stay; P(h): how likely each one is; B(h): nights actually booked in it. E[B] is the expected number of booked nights.

No state-space model needed. You get an estimate with its uncertainty attached. You can also use the minimum stay as a probe: a gap exactly as long as the minimum stay that closes is almost certainly a real booking. A longer gap carries a known, bounded risk: at most m − 1 stranded nights on each side of each stay.

Which Airbnb ADR and occupancy figures you can trust

Not every figure is equally wrong. The error depends on whether a booking can leave a gap shorter than the minimum stay, and that comes down to the listing's rules. The table uses the same one-bedroom apartment at about 65% occupancy, with only the minimum stay changed.

Revenue overstated by calendar-based data, by minimum stay. Simulated example, one year.
Minimum stayCan nights be sandwiched?OverstatedHow far to trust it
1 nightNo. Any open night can be booked.0%Clean
2 nightsSingle nights between staysabout 6%Good
3 nightsUp to 2 nights on each side of a stayabout 10%With caution
5 to 7 nightsUp to 4 to 6 on each sideabout 16%Upper bound only
14 to 30 nights, any check-in dayUp to 13 to 29 on each sideabout 18 to 19%Least reliable
30 nights, whole calendar months onlyNo, apart from Februaryclose to 0%Clean

Two other things move a listing along this scale. Busier calendars break into more short gaps, so a popular listing sits further toward the unreliable end than a quiet one with the same rules. And check-in day restrictions help: the more a host lines stays up with the minimum, the fewer leftovers there are.

What this means for Airbnb investors and hosts

  • Treat calendar-based revenue as a ceiling, not a forecast. Especially in markets with long minimum stays, and especially for top-performing comps.
  • Check the comps' minimum stays. Comps with 1-night minimums or whole-month rules give you clean numbers. Comps with 3 nights or more, and any check-in day, give you a ceiling.
  • In New York or any 30-day market, discount hard, unless the listing only takes whole calendar months. Otherwise a single booking can move its apparent occupancy by a full month.
  • Ask your data provider how they handle gaps below the minimum stay. If the answer is "our model classifies them", ask how it can, when the calendar is identical either way.
  • Watch your own orphans. Gap-night discounts or a lower minimum for short gaps turn dead nights back into bookable ones.

So what can you trust?

Calendar-based Airbnb data isn't wrong everywhere. It's wrong in predictable places. Where no booking can leave a gap shorter than the minimum stay, the calendar tells the truth: listings with a 1-night minimum, and 30-night listings that only take whole calendar months. Build your comparison set from those, and their occupancy, ADR and revenue figures are as close to real as scraped data gets.

Everywhere else, the figures lean one way, too high, and they lean further the longer the minimum stay and the busier the listing. Use them as an upper bound, check them against the clean comps, and ask your data vendor how it handles nights sandwiched by minimum stays. The data can still guide a decision. You just need to know which numbers are carrying the weight.

Frequently asked questions

How accurate is Airbnb occupancy and ADR data?

Scraped Airbnb data counts a night as booked when it turns unavailable on the public calendar. Nights sandwiched between stays by the minimum-stay rule close without being booked, so occupancy and revenue come out too high. The error is small with 1- or 2-night minimums and grows with longer minimum stays and busier calendars.

Which Airbnb occupancy and ADR figures are most reliable?

Figures for listings where no booking can leave a gap shorter than the minimum stay: listings with a 1-night minimum, and 30-night listings that only accept whole calendar months. Busy listings with minimum stays of 3 nights or more and any check-in day are the least reliable.

What is an orphan night on Airbnb?

A free night, or a short run of free nights, left between two bookings that is shorter than the listing's minimum stay. No guest can book it, so it earns nothing, but it's often closed on the calendar by the host, a channel manager or a gap rule.

Why does the minimum night stay affect Airbnb occupancy estimates?

The higher the minimum stay, the more nights a single booking can strand. With a 2-night minimum a booking can strand 1 night on each side. With a 30-night minimum, as in New York, up to 29 on each side. Each stranded night that closes looks like a booked night to anyone reading the calendar from outside.

Are busy listings overestimated more than quiet ones?

Yes. Every booking is another chance to strand nights, and busy calendars are split into short gaps where orphans form most easily. In our simulation with a 3-night minimum, a listing with about 9 stays a year had 1% phantom nights, while one with about 62 stays had over 10%.

Can machine learning tell booked nights from blocked nights?

Not from public calendars alone. A booked night and a closed orphan night leave the same trace, so the information needed to tell them apart isn't in the data. A model can estimate how often it happens and discount for it, but it can't recover which nights were booked.

Sources

  1. [1]
    NYC Mayor's Office of Special Enforcement, Short-Term Rental Registration (Local Law 18)