charlesreid1.com blog

Fermi Problems Are Chocolate Messes All the Way Down

Posted in Mathematics

permalink

In a prior post Fermi Problems are Quadrature Problems we worked through the classic Fermi problem "how many chocolate bars are eaten each year in the United States?" twice - once as a naive one-point estimate (~40 billion bars/year) and once as a refined 4×4 age-by-season quadrature that preserved the age/season interaction (~44 billion bars/year). Both estimates hovered in the same order of magnitude, which was satisfying but also unfalsified: we worked from first principles and back-of-the-envelope estimates. No published, macroeconomic data from the real world went into the estimates.

This post is the falsifiability check. It turned into a blog post because it turned out to be much more interesting than "look up two or three published numbers and back out the real answer" because "real" is hard to pin down, there isn't actually an "answer," and the rabbit hole keeps going deeper, because building out bounds for the original Fermi problem is itself a Fermi problem. And the numbers that go into those Fermi problems pack in many assumptions that can themselves be unpacked into Fermi problems.

It's Fermi problems all the way down.

The Aggregator Trap

The obvious first move is to do a web search for "how much chocolate do Americans eat per year" and take the top result. Do that (like a rube) and you get some variation of these two claims, repeated across dozens of sites:

  1. Americans eat 2.8 billion pounds of chocolate per year (~11 lb/person).
  2. The average American eats about 3 chocolate bars per week.

Both numbers convert cleanly into "bars per year." Claim 1, at a standard 1.55 oz Hershey's bar (~10 bars/lb), gives 28 billion bars/year. Claim 2, over 52 weeks × 340M people, gives 53 billion bars/year. That would be a tidy [28B, 53B] corridor that both of our estimates sit inside. Post over. Ship it. Except that it is total bullsh-t.

Sites using this claim are a giant Gordian knot of self-referential listicle sites, paid aggregators, and misspelled/garbage URLs grifting as "legitimate" sources.

Quality score: 1/10.

The Actual Primaries

Two sources survive scrutiny:

USDA Economic Research Service publishes cocoa bean import volumes as part of its agricultural trade tracking. Between 2000 and 2022, the US imported an average of ~425,000 metric tons of cocoa beans per year. (2023 and 2024 were unusual crop-failure years and dropped to 269kt and 198kt respectively; we'll use the long-run average since the recent dip is a supply shock, not a demand signal.) This is a real customs-derived number with a known methodology.

(Side note: reporing an average consumption per year over TWENTY YEARS reeks of Fermi sub-problems.)

National Confectioners Association's State of Treating 2025/2026 report, compiled using Circana retail-panel data and Euromonitor market modeling, reports US chocolate sales of $28.4 billion in 2025 (51.7% of the $55B total confectionery market). This is a real industry number with named data providers.

To turn either one into "bars per year," we have to introduce multiple layers of estimates (models). Which is to say, we have to do a Fermi problem.

Lower-Bound Fermi Problem: Cocoa Tonnage to Bars

Starting from 425,000 metric tons/year of imported cocoa beans, the chain of unit conversions goes something like:

  1. Bean → cocoa mass. After shell removal, roasting, and winnowing, yield is roughly 80% of raw bean mass. ~340,000 t of usable cocoa mass.
  2. Cocoa mass → finished chocolate. Finished chocolate is a blend. Milk chocolate is ~10–20% cocoa content by mass; dark is 50–85%; white is 0% solids. Weighted by US market mix (heavily milk), call the average cocoa content ~20% → 340,000 / 0.20 → 1.7 million metric tons of finished chocolate → 3.75 billion pounds.
  3. Non-bar uses. Some cocoa becomes baking cocoa, hot chocolate mix, cosmetic cocoa butter, ice cream inclusions - not bars. Call it 15–25% of cocoa mass. Subtract ~20%.
  4. Net trade in finished chocolate. The US is a net importer of finished chocolate (Toblerone, Lindt, Cadbury via Canada, etc.), which adds back roughly the same order of magnitude as the non-bar leakage. Call these a wash to first order.
  5. Pounds → bars. At 1.55 oz/bar → 10.3 bars/lb → 3.75 × 1e9 lb × 10.3 → ~39 billion bars/year.

Every step in that five-step decomposition carries real uncertainty. Bean yield could be 75% or 85% (±6%). Average cocoa content could be 15% or 25% (±25%). Non-bar share could be 10% or 30% (±33%). Net trade is a wash only to first order. Individually these are tens-of-percent knobs, not factor-of-two ones - but they compose multiplicatively, and five stacked ±25%-ish knobs is already enough to smear the answer by roughly ±2×. And every arrow in the chain is a Fermi sub-problem whose inputs are numbers someone else estimated the same way.

If we swing the knobs pessimistically, we land at ~25B bars. Optimistically, ~55B. The "lower bound from real primary data" is a range of ~25–55B, which is not really a bound at all - it's a Fermi estimate wearing a lab coat.

Upper-Bound Fermi Problem: Dollar Sales → Bars

Starting from $28.4B/year in US chocolate sales, the chain is shorter but each step is fuzzier:

  1. Dollar sales → bar-equivalents. We need an average retail price per "bar." Full-size checkout bar: ~$1.50–$2.00. Premium bar (Lindt, Ghirardelli): $3–$5. Fun-size piece from a Halloween bag: ~$0.15–$0.30. Bulk seasonal chocolate (Easter eggs, Valentine's boxes, advent calendars): wildly varying $/oz, but each piece often counts as an eating event.
  2. Weighted average. If bar-form full-size chocolate dominates sales dollars, the average is close to $1.50/bar. If fun-size and bulk dominate the count, the average is closer to $0.50/piece.

That's not a factor of two; it's a factor of three, and the whole answer sits on it:

Average $/bar Implied bars/year
$0.50 ~57 billion
$1.00 ~28 billion
$1.50 ~19 billion

The "upper bound from real primary data" is anywhere from 19B to 57B, depending entirely on how we resolve one knob - which is itself the same "what counts as a bar" linguistic knob we called out in the very first section of the prior post. Primary data didn't retire the knob. It just relocated it into the unit-conversion step.

Fermi Problems All the Way Down

Here's the pattern. We started with a Fermi problem (bars/year). To bound it, we went hunting for real data. The real data we found isn't denominated in the units we want, so converting it into bars/year required building another Fermi problem - a decomposition into estimated factors each guessed to within a factor of two. And when we look at where those factors come from, they are the outputs of other Fermi problems that somebody else already solved:

  • USDA's 425,000 metric tons of cocoa imports is itself the output of a chain: customs declarations aggregated across ports, unit conversions, imputations for missing data, seasonal adjustments. USDA analysts did their own Fermi decomposition to produce that number; we just don't see the pieces.
  • NCA's $28.4B in chocolate sales is Circana's retail-panel extrapolation (some measured stores × a scaling factor for coverage) blended with Euromonitor's channel-mix model. Both of those are big decomposition chains ending in a single reported number.

There is no bedrock number that lives outside a Fermi decomposition. Every "primary" figure is the leaf node where someone earlier in the chain decided to stop decomposing and report a total. Push on any leaf hard enough and you find yet another decomposition underneath.

And the decomposition doesn't just go deeper; it goes sideways into whole other disciplines. Take the "$/bar" knob from the upper-bound problem. We handwaved it as $0.50 to $1.50 depending on whether fun-size or full-size dominates. But the honest way to attack that knob is to estimate the joint distribution of chocolate bar types across the US market - Hershey's checkout bars vs. Costco Toblerones vs. Halloween fun-size vs. seasonal boxed chocolates vs. premium Lindt vs. bulk baking chocolate - and that is an entirely different subclass of Fermi problem. It lives in economics and finance, not in demography. It asks:

  • Which manufacturers dominate US chocolate volume? (Hershey, Mars, Ferrero, Mondelez, Lindt, Nestlé, Ghirardelli, and a long tail.)
  • What is each manufacturer's product mix, and what fraction of their US revenue is bar-form vs. boxed vs. seasonal vs. inclusions?
  • What are their per-unit margins, wholesale prices, and retail markups?
  • What's the channel distribution - grocery vs. drug vs. mass vs. club vs. convenience vs. specialty - and how does the price-per-bar-equivalent vary across those channels?

Each of those questions is itself a Fermi problem, and now the primary data isn't customs tonnage or retail-panel dollar totals - it's public companies' 10-K filings, segment reporting, investor decks, and gross margin disclosures, and you're reverse-engineering unit economics out of consolidated financial statements that were never designed to answer "how many bars." Hershey's 10-K tells you global net sales for the "North America Confectionery" segment, but not how many kilograms it shipped, and definitely not how many bars; you have to estimate the mapping using industry-average margin data, competitor benchmarks, and public price-per-unit surveys - each of which is another Fermi sub-problem several layers down.

So the recursion isn't a single tree of decompositions; it's a forest. Different branches recruit different fields - demography for the population count, agronomy for cocoa yield, industrial chemistry for processing losses, corporate finance for the manufacturer-and-margin map. Each field has its own conventions for what counts as a "primary" number and where the field stops decomposing. Cross-field, those conventions don't line up, so when you try to reconcile a demographic estimate with a financial one you're not just multiplying uncertainties - you're translating between whole vocabularies of what "measured" means.

Which means our two bounds - the ~25–55B range from cocoa tonnage and the ~19–57B range from dollar sales - are not really bracketing the truth so much as recording where two different chains of Fermi decompositions happened to land. Union those ranges and we get roughly [19B, 57B], which comfortably contains our two original estimates (40B and 44B). But we should be honest about what that means: it doesn't mean the original estimates were validated. It means every path we can take to an answer inherits the same tens-of-percent-per-knob budget, and those knobs compose multiplicatively, so a five-step chain of ±25%-ish knobs smears the answer by roughly ±2× even if every individual knob is honest - and if any single knob is genuinely factor-of-two (like the $/bar knob in the upper-bound problem), it dominates the whole budget by itself.

The interesting thing isn't that our Fermi estimate landed in the middle of the corridor. The interesting thing is that the corridor is the same width as the original problem. We didn't tighten the answer by bringing in real data. We just relocated the uncertainty from "what's the per-capita rate?" into "what fraction of cocoa becomes bars?" and "what's the average price per bar?"

The Fractal Coastline

In a prior post Fermi Problems are Quadrature Problems took a one-sentence bar-trivia question and pulled on the thread until it turned into the population balance equation - the same formalism used to track particles in a turbulent jet, or precipitation of ice particles in the atmosphere contributing to cloud seeding and growth - but applied to Halloween candy and demographic distributions.

This post pulls on the next thread. The moment you try to bound the answer against reality, the "reality" you were reaching for turns out to be another Fermi problem, and each of its inputs is another Fermi problem, and the branches recruit different fields - demography, agronomy, industrial chemistry, corporate finance - each of which terminates its own decomposition at whatever level of aggregation happened to be convenient for its own purposes. There is no floor. The "primary data" you were hoping to stand on is a platform someone else built by decomposing their problem far enough to feel done, and if you step on it and push, it decomposes further.

This is not a bug in Fermi estimation. It's a feature of the world. Reality is a fractal coastline: the closer you look, the more structure appears, and the total "length" you measure depends entirely on the size of the ruler you brought. Every time we ask "how much chocolate?" we're really asking "at what altitude?" A satellite photo says one thing; a customs manifest says another; a Hershey 10-K says a third; a grocery store receipt says a fourth; a five-year-old's Halloween pillowcase says a fifth. None of them are wrong. They are measurements at different scales, and the scales don't compose into a single tidy total because the coastline doesn't have a single tidy length.

What the Fermi decomposition buys us is not a number. It's a map of the coastline at the altitude we chose to fly at, with every landmark labeled: here is where we assumed a population, here is where we assumed a rate, here is where we assumed a cocoa content, here is where we assumed a price per bar. Every landmark is a place we could descend and zoom in further, and the map tells us exactly what we would gain (and lose) by doing so. That's not a consolation prize for failing to find the "real" answer. That is the real answer, at the altitude the question was asked.

The aggregator's confident 2.8 billion pounds is a photograph of the coastline with no scale bar. Our 40 billion bars is a photograph with the scale bar drawn on it, plus a legend explaining what every marking means and how the picture would change if you flew lower. The picture isn't sharper - it's actually a little fuzzier, because we've been honest about the fuzz - but for the first time it's a picture you can navigate by.

And that's the whole thing. We aren't strolling along on solid ground counting Snickers wrappers. We're traversing a fractal coastline, one moment in time at a time, with a map we drew ourselves, at whatever altitude we chose.

It is a chocolate mess. But it is a beautiful chocolate mess.

References

Tags:    mathematics    fermi problems    chocolate   

Fermi Problems Are Quadrature Problems

Posted in Mathematics

permalink

Here's a Fermi problem you'll run into sooner or later, usually in a job interview, sometimes over drinks with someone who wants to see how you think:

How many chocolate bars are eaten each year in the United States?

The provenance goes back to Enrico Fermi's Chicago physics classes, where he'd ask students to estimate the number of piano tuners in the city, or the width of a nail's head in miles, or anything else where the point wasn't the number but the decomposition. That habit of mind has since been laundered through McKinsey and Google into the standard interview format, but the underlying move is Fermi's: break a hard question into smaller questions you can each guess within a factor of two, and let the errors mostly cancel.

Does Size Matter?

Before we start estimating, there's an elephant in the room that deserves to be named and then, deliberately, ignored. What counts as "a chocolate bar"? A fun-size Snickers you get in a Halloween bucket? A king-size Twix? A square of a Hershey's bar? A Costco-sized Toblerone? The answer changes the final count by an arbitrary factor of 1/2 or 2 or worse - it's the single variable this whole problem is most sensitive to, and it's not really a quantitative question at all. It's a linguistic knob. Turning it doesn't teach us anything about populations or quadrature; it just relabels what we're counting.

So we're going to do what physicists do when a problem has a knob like this: we're going to assume a spherical cow. In the spirit of Fermi, "eating a chocolate bar" is an idealized, dimensionless event - a discrete unit of chocolate consumption, roughly the size of whatever the reader pictures when they hear the phrase. We're not going to quibble about grams or servings or whether a Reese's cup counts. If your definition differs from mine by a factor of two, your final answer differs from mine by a factor of two, and that's fine. The interesting structure of this problem lives in the population and its interactions, not in the definition of the counting unit, and the machinery we're about to build works the same way whichever definition you pick.

With that noted and set aside, on to the actual estimation.

The Naive Approach and What It's Secretly Assuming

The textbook approach to the chocolate question looks like this. Guess that the average American eats maybe one chocolate bar every few days - call it 0.3 per day. Multiply by the population, roughly 340 million. Multiply by 365. You get something on the order of 40 billion chocolate bars per year.

That's your answer, and if you've picked reasonable numbers it's probably within a factor of two or three of the truth.

But look at what that multiplication is quietly assuming: one average American, one average day, one average rate. Every source of variation in the real population has been smoothed into a single number. If you push on the estimate, the assumption starts to feel thin. A five-year-old on Halloween is not eating the same amount of chocolate as an fifty-year-old on Halloween. Neither of them is eating the same amount on Halloween as they are in March. And it gets worse when you notice that the seasonal spike itself depends on who you are: kids own Halloween, adults own Valentine's Day, families with young children own Easter. (We can pretty safely assume that seniors will be flat across all of it.)

You can try to refine the estimate by splitting into age buckets, or by splitting into seasonal buckets. Either helps a little. But do them separately and you still miss the thing that actually matters, which is that age and season interact. A "seasonal multiplier" averaged over the whole population says everyone eats three times more chocolate on Halloween. That's false in a specific and important way: kids eat ten times more, adults eat only slightly more, and the average is a fiction that lives in neither group.

The interaction is real, and no amount of separately-refined marginal averages will recover it. We need a formalism that lets us write the joint structure down directly.

Population as a Joint Density (or, A Little Quadrature Never Hurt Anyone)

Here's the reframe. Think of the population as a density over some space of coordinates. Each person has internal coordinates \(\xi\) - age, income, dietary preferences, whatever matters for the problem - and lives in the external coordinate of time \(t\). Let \(n(\xi, t)\) be the number density of people: \(n(\xi, t)\, d\xi\) is how many people have coordinates in a little box around \(\xi\) at time \(t\). Let \(r(\xi, t)\) be the per-capita consumption rate: chocolate bars per person per unit time, for a person with coordinates \(\xi\) at time \(t\).

Then the total number of chocolate bars eaten in a year is just the integral of the product:

$$ \text{CBE} = \int_0^{1\,\text{yr}} \int_\xi r(\xi, t) \, n(\xi, t) \, d\xi \, dt $$

Two objects, cleanly separated. \(n\) says who exists. \(r\) says what they do. The product \(r \cdot n\) is chocolate bars per unit coordinate per unit time, and integrating it over the whole domain gives the total. (This is the population balance equation, a workhorse in chemical engineering, particle dynamics, and demography. We're borrowing the machinery, not reinventing it.)

Now the key observation: any numerical evaluation of that integral is a quadrature. Pick a grid of abscissae \((\xi_i, t_j)\), assign each grid cell a weight \(w_{ij}\) (in the simplest case, just the cell's width), and sum:

$$ \text{CBE} \approx \sum_{i,j} w_{ij} \, r(\xi_i, t_j) \, n(\xi_i, t_j) $$

That's the whole game. The naive Fermi estimate we started with - rate × population × 365 - is exactly this sum with a single term: one \(\xi\)-bin covering everyone, one \(t\)-bin covering the whole year, one average rate. It's a one-point quadrature over the joint density. Refining a Fermi estimate is refining a quadrature grid. Everything from the back-of-envelope guess up to a full numerical integration lives on the same continuum.

And crucially, this formulation does not require \(r\) to factor as an age-dependent piece times a seasonal piece. You fill in each grid cell independently, so any interaction between coordinates is preserved by construction. That's exactly the structure the marginal-average approach destroys.

There's one more idea worth borrowing from numerical analysis here. The whole art of Gaussian quadrature is that abscissae should not be evenly spaced - they should cluster where the integrand varies fastest. The same instinct applies to Fermi estimation: a good quadrature grid matches the shape of what you're integrating, not the shape of the axis you're integrating over. For the chocolate problem, the "time" axis has twelve months, but the integrand has three spikes: Halloween, Valentine's, Easter. Using twelve evenly-spaced monthly bins would spend most of your resolution on quiet months where the integrand is flat and boring. Concentrating your grid on the three candy-heavy windows plus one bin for "everything else" captures almost all the structure with a quarter of the work. That reduction - twelve bins to four - is not a shortcut; it's the correct grid for this integrand.

There's a subtlety worth naming here, though. When we collapse twelve evenly-spaced months into four unevenly-sized "seasons," we're doing something a little sneaky: we're letting the content of the year (which holiday, which age group's spike) reshape the time axis itself. What was a pure external dimension - clock time, marching uniformly forward - is being warped into an internal dimension that carries information about the integrand. That works cleanly here because there's no other process on the time axis we care about. As problems gain dimensions, though, that alignment can break.

(As a more involved example, if we were trying to approximate number of chocolate bars eaten by a population, but under conditions of continuous changes in population or demographics, then integrating the total population over time would need a properly-resolved time axis - lumping 46 weeks into one bin would be throwing away real information. The seasonal grid works for our problem, because the integrand's structure and external axis's structure happen to coincide. If they don't, keep them separate, and "pay" for the extra abscissas (with a bit more bookkeeping).

Working the Chocolate Problem With Interactions

With the framework in hand, the estimate becomes a table. Rows are age buckets, columns are seasonal windows, each cell holds the number of bars eaten by that group in that window.

For the age axis, four buckets are enough: kids (0-12), teens (13-19), adults (20-64), and seniors (65+). Population is roughly uniform over these bands, so with 340 million Americans we can allocate roughly 55M kids, 30M teens, 190M adults, and 65M seniors. (Uniform-in-age is a lie, but a small enough one that the drama of this problem lives in \(r\), not in \(n\). That itself is a useful diagnostic - it tells us where refining would and wouldn't help.)

For the time axis, four bins: a Halloween window (~2 weeks around Oct 31), a Valentine's window (~2 weeks around Feb 14), an Easter window (~2 weeks around the spring holiday), and the remaining ~46 weeks of the year lumped into "rest of year." Numbers below are bars per person per day, and we'll multiply through by bucket population and bin length at the end.

Halloween (14d) Valentine's (14d) Easter (14d) Rest (312d)
Kids (55M) 3.0 0.4 1.5 0.3
Teens (30M) 2.0 0.6 0.4 0.4
Adults (190M) 0.5 1.5 0.3 0.25
Seniors (65M) 0.2 0.3 0.2 0.15

Look at that matrix for a second. The kids row peaks on Halloween. The adults row peaks on Valentine's. Easter has a bump for kids and almost nothing for anyone else. Seniors are flat and low. No product of a row-vector and a column-vector produces this pattern - the matrix is not rank-1, and any factorization \(r(\xi, t) = r_{\text{age}}(\xi) \cdot s(t)\) would flatten these ridges into a smooth surface that gets every cell wrong. This is precisely the structure the joint formulation preserves and the marginal-refinement approach loses.

Summing (population × rate × days) cell by cell:

$$ \text{CBE} \approx \sum_{i,j} n_i \cdot r_{ij} \cdot w_j \approx 44 \text{ billion bars/year} $$

Which is, satisfyingly, in the same order of magnitude as the naive estimate from the first section (40 billion for the naive approach, 44 billion for the quadrature estimate). That's the usual outcome, and it's not a knock on the framework - it's a reminder that averaging over correlated variables often lands close to the truth by luck. What the framework buys you isn't necessarily a better number; it's a legible number. Every approximation is named and located in a specific cell. If you wanted to sharpen the estimate, you'd know exactly where to add resolution - split Halloween into "trick-or-treat night" and "the week after," split kids into "young enough to trick-or-treat" and "too cool for it" - and you'd be adding grid points to the ridges, which is exactly what a higher-order quadrature scheme does automatically.

That's the transferable idea, and it generalizes beyond chocolate. Any Fermi problem about a population - how many haircuts per year, how many gallons of coffee consumed per week, how many miles driven per day - is an integral of a per-capita rate against a number density over some joint coordinate space. The reason the standard "multiply the averages" trick works at all is that most people's mental \(r\) is smooth enough that a one-point quadrature suffices. When it isn't - when the coordinates interact, when the ridges matter - the framework tells you exactly where to add resolution and why. You stop guessing a single average and start choosing a grid.

References

Tags:    mathematics    fermi problems    quadrature    numerical methods    chocolate   

Nmap Host Discovery: All the Ways to Ask "Is Anyone There?"

Posted in Security

permalink

This is a companion post to Building an Nmap Short Course from Scratch. Where that post was about the meta - course design, lab infrastructure - this one drills into the actual first-lecture material: how Nmap decides whether a host is up.

Full lecture notes: Nmap/Short Course/Lecture 1.

Why Host Discovery Matters

Before you can scan ports, identify services, or check for vulnerabilities, you have to figure out which IP addresses on the target network actually have a machine behind them. Scanning IPs that aren't responding is a waste of time, generates a lot of unnecessary network noise, and can tip off defenders.

Think of it as making a map of active settlements before deciding which ones to explore in detail.

Ethics First

The obvious but necessary caveat: Nmap must only be used on networks where you have explicit, written authorization to scan. Unauthorized scanning can be interpreted as an attack, and in many jurisdictions is illegal.

Everything in this post assumes an isolated lab environment - which is exactly what we set up in the previous post.

The Default: Nmap's Multi-Probe Approach

If you run Nmap as a privileged user (root or sudo) without specifying any discovery options, it will fire four probes at each target:

  1. ICMP echo request (a classic ping)
  2. TCP SYN packet to port 443
  3. TCP ACK packet to port 80
  4. ICMP timestamp request

If any of the four gets a response, Nmap considers the host up.

The multi-probe approach exists because different firewalls block different things. ICMP is commonly blocked at the network edge. Port 443 might be allowed inbound because there is a web server behind it. Port 80 might respond with a TCP RST because there is nothing listening. Any one of these signals is enough.

Unprivileged users can't send raw packets, so Nmap falls back to attempting TCP connect() calls to ports 80 and 443. Less accurate, but works without root.

-sn: Just Tell Me What's Alive

Nmap's default behavior is to do host discovery and then port scan whatever comes back alive. If you only want the host discovery step, use -sn:

nmap -sn 192.168.1.0/24

-sn means "scan, no port scan." (In older versions it was -sP, "scan ping.") It runs the multi-probe discovery and prints just the list of live hosts. Very fast, very quiet compared to a full port scan, and often the first thing you run.

Everything else in this post uses -sn unless otherwise noted.

The -P Family: Picking a Specific Probe

If you want to control exactly which probe Nmap sends, use one of the -P flags.

-PE: ICMP Echo (the plain ping)

sudo nmap -sn -PE 192.168.1.100

The most familiar probe. Sends an ICMP echo request, expects an ICMP echo reply. Works when firewalls allow ICMP. Frequently blocked at network perimeters.

-PP: ICMP Timestamp

sudo nmap -sn -PP 192.168.1.101

Sends an ICMP timestamp request (type 13), expects a timestamp reply (type 14). Useful when echo requests are blocked but timestamp requests aren't - some firewall rules block ICMP type 8 (echo) but forget about type 13.

-PM: ICMP Address Mask

sudo nmap -sn -PM 192.168.1.102

Sends an ICMP address mask request. Very rarely used legitimately these days, which is exactly why it sometimes gets through firewalls that block the more common ICMP types.

The pattern: try the obvious probe, fall back to the less-obvious ones if the obvious one fails.

-PS[ports]: TCP SYN Ping

sudo nmap -sn -PS 192.168.1.0/24
sudo nmap -sn -PS22,80,443 192.168.1.50

Sends a TCP SYN packet to the given ports (default 80 if you don't specify). A response - either SYN/ACK meaning "port open" or RST meaning "port closed" - tells Nmap the host is alive.

This one is the workhorse against firewalled targets. Firewalls typically allow inbound traffic to common service ports (80, 443, 22) because there are legitimate reasons for outsiders to reach those ports. TCP SYN pings ride on that permitted traffic.

-PA[ports]: TCP ACK Ping

sudo nmap -sn -PA 192.168.1.0/24
sudo nmap -sn -PA21 192.168.1.55

Sends a TCP packet with the ACK flag set. This is a weird packet - an ACK with no prior SYN - so most operating systems respond with a TCP RST regardless of whether the port is open. If you see the RST, the host is alive.

Useful against stateful firewalls that block unsolicited SYN packets (because they aren't part of any tracked connection) but let ACKs through (because ACKs look like the middle of an established connection the firewall might have lost track of).

-PU[ports]: UDP Ping

sudo nmap -sn -PU 192.168.1.0/24
sudo nmap -sn -PU53,161 192.168.1.60

Sends a UDP packet to the given ports (default 40125, chosen because it's usually closed). If the port is closed, the host should respond with ICMP port unreachable, telling you it's alive. If the port is open, you might not get a response at all.

UDP ping is less reliable for host discovery on its own, but very useful when the target runs UDP services (DNS on 53, SNMP on 161) and when other probes are all blocked.

-PR: ARP Ping

sudo nmap -sn -PR 192.168.1.0/24

The gold standard when you are on the same Ethernet segment as your targets. ARP is the layer-2 protocol that resolves IP addresses to MAC addresses, and hosts cannot refuse to answer ARP - if they did, they would be unable to talk to anything on the local network.

Nmap automatically uses -PR for local-segment targets when run by a privileged user, unless you tell it not to (--send-ip). It's fast (no round trip past the switch) and 100% reliable for hosts that are up.

Target Specification

Independent of the probe type, you need to tell Nmap which IPs to scan.

Single addresses

nmap 192.168.1.1
nmap scanme.nmap.org

Hostnames get resolved via DNS. scanme.nmap.org is a target the Nmap project maintains specifically for people to practice against.

CIDR ranges

nmap -sn 192.168.1.0/24
nmap -sn 10.0.0.0/8

Standard CIDR notation. /24 is 256 IPs, /8 is 16.7 million (be careful).

Numeric ranges and lists

nmap -sn 192.168.1.1-100        # .1 through .100
nmap -sn 192.168.1.1,2,10,50    # specific IPs
nmap -sn 192.168.1,2,3.1-254    # cross product

That last one scans .1-.254 for each of 192.168.1.x, 192.168.2.x, and 192.168.3.x. Useful for scanning a handful of adjacent subnets.

From a file

nmap -sn -iL targets.txt

One target per line in the file. Best option when you have a large or irregular list of targets, or when the target list comes from another tool.

Excluding targets

nmap -sn 192.168.1.0/24 --exclude 192.168.1.1,192.168.1.100
nmap -sn 192.168.1.0/24 --exclude-file dontscan.txt

Critical for avoiding accidental scans on the CEO's laptop or the production database when you're supposed to be scanning a specific subnet.

Timing: -T0 through -T5

Nmap has six timing templates that control how aggressively it sends packets:

  • -T0 (paranoid): Extremely slow. One probe every few minutes. Used for IDS evasion in serious red-team engagements.
  • -T1 (sneaky): Slow. IDS evasion, but not as extreme.
  • -T2 (polite): Slower than default. Reduces bandwidth and target load. Good for scanning production infrastructure.
  • -T3 (normal): The default. Reasonable timing for most networks.
  • -T4 (aggressive): Faster. Assumes a reliable network. Good balance for lab work.
  • -T5 (insane): Very fast. Sacrifices accuracy for speed. Can overwhelm slow networks or fragile targets.
sudo nmap -sn -T4 192.168.1.0/24

For host discovery specifically, -T4 is usually the right choice in a lab environment. On production, use -T3 or -T2 and be patient.

Reading the Output

A successful scan looks like this:

Starting Nmap 7.94 ( https://nmap.org ) at 2025-05-27 14:00 PDT
Host 192.168.1.1 is up (0.00050s latency).
MAC Address: AA:BB:CC:DD:EE:FF (Realtek Semiconductor)
Host 192.168.1.10 is up (0.00080s latency).
MAC Address: 11:22:33:44:55:66 (VMware)
Nmap done: 256 IP addresses (2 hosts up) scanned in 2.10 seconds

The MAC address only shows up when you're on the same Ethernet segment. The vendor in parentheses comes from Nmap's built-in OUI database - useful for spotting VMs, or figuring out which switch port a device is behind.

When Hosts Don't Appear

"Host is down" in Nmap's output really means "Nmap didn't get a response from any of the probes it sent." That's not the same as "the host is offline." Common reasons a live host doesn't appear:

  • Restrictive firewall. The most common culprit. Dropping all probe types silently is a valid (if aggressive) defensive posture.
  • Host-based firewall. Windows Firewall, iptables, and similar can block probes even when the network firewall is permissive.
  • Wrong scan for the environment. ICMP-only discovery against a target that only accepts TCP-80 - the host won't appear even though a -PS80 scan would find it in an instant.
  • Unprivileged Nmap. Non-root Nmap has very limited discovery options. Always run as root (or via sudo) for real scanning.
  • Network layer issues. Routing problems, wrong subnet mask on the scanner, VLAN mismatches - anything that prevents packets from reaching the target.

The right response to a "no hosts up" result on a network you know is alive: try a different discovery method. If ICMP fails, try -PS22,80,443. If TCP fails, try -PA. If nothing works, you're probably up against a very well-configured firewall.

From Discovery to Deeper Scans

Once you have a list of live hosts, the natural next step is port scanning them to see what services are running. The clean way to chain this is with grepable output:

nmap -sn -oG - 192.168.1.0/24 | awk '/Up$/{print $2}' > live_hosts.txt
nmap -sV -iL live_hosts.txt

The first command produces a machine-parseable list of live IPs. The second feeds that list into a service-detection scan. This two-phase approach is efficient (you don't waste time port-scanning dead IPs) and organized (you have a saved list of live hosts to work from later).

Service detection, port scanning, NSE scripting - all of that comes in later lectures of the course. But it all starts with knowing who's home.

References

Tags:    security    nmap    host discovery    ping    arp    networking    pentesting   

March 2022

How to Read Ulysses

July 2020

Applied Gitflow

September 2019

Mocking AWS in Unit Tests

May 2018

Current Projects

November 2017

A Hard(y) Math Problem