First contact

The visitor knocked quietly before it walked in. Its first two requests came on 16 August from a single address, 47.79.233.158, asking only for / and /robots.txt. A month later, on 16 September, 47.82.50.25 asked for /cfdemo.html.
On 18 September it began probing in earnest from many addresses in 47.82.0.0/16, asking for the things every secret-hunter wants: /phpinfo.php, /uploads/export_users.sql, /aws-credentials. The MIRE answered them all with bait. Somewhere in that sweep the crawler requested /dev, /test or /staging, and found a page that looked like a server left wide open.
The endless corridor

The fake staging page was built as a maze on purpose. Each page carries three links built from its own address, to keep a curious bot clicking a little longer. Nobody expected anything to follow them forever:
<a href="{request.path}/status.php">System Status</a>
<a href="{request.path}/health">Health Check</a>
<a href="{request.path}/config">Configuration</a>
# what the crawler was asking for by 9 October
GET /dev/config/status.php/health/config/status.php/config/config/health/config/config
So /dev linked to /dev/config, which linked to /dev/config/health, and so on with no bottom. Each page opens three new doors, so level 10 alone holds 310 = 59’049 pages, and every one is served with fresh fake database passwords and API keys. By 9 October over half the crawler’s requests were more than 10 levels deep.
The crawler worked level by level, and each level is three times wider than the last. Its daily total stepped up roughly every second day, from 299 on 28 September to 41’233 on 9 October. Across all MIRE hosts, logged requests went from about 10’000 a day to more than 54’000.
Why nothing slowed it down

The MIRE already punishes repeat visitors, but this crawler never looked like one. The built-in delay grows by 0.1 s per request from the same IP, up to 5 s. This crawler spread its work so thin that no single address ever built up a penalty.
| Trait | What we saw (4–9 October) |
|---|---|
| Addresses | 4’004 unique IPs, 99% in 47.82.0.0/16 and 47.79.0.0/16 (Alibaba Cloud) |
| Location | Cloudflare country SG on 108’586 of about 110’000 requests |
| User agent | Desktop Chrome 144–150 on Windows and macOS, each version about 5’500 times |
| Headers | Full browser set: Sec-Ch-Ua, Sec-Fetch-*, Accept-Language: en-US |
| Targets | /dev 40’082 · /test 37’273 · /staging 31’233 |
| Other sites | A few hundred hits on mire.cc, toce.ch, willigetpwned.com; almost all effort went into the maze |
It looked like a fleet of ordinary browsers, which is exactly how a well-funded scraper tries to look. The one thing it could not hide was where it went: paths ten levels deep that no person would ever type.
Keep it stuck, but not forever

We chose to keep the crawler in the mire but make every step cost it time, and to give the maze a floor. Leaving it alone was not an option: the real danger was The MIRE taking itself down because the trap was working.
- Logs: the cfd.mire.cc Caddy log rolled over at 150 MB twice in three days. The MIRE’s access.log reached 850 MB and useragent.log 350 MB.
- Noise: 104’000 of about 140’000 MIRE requests since 5 October came from this one crawler, burying the probes we actually want to study.
- Growth: at a doubling every two days, it would pass 100’000 requests a day within a week. Every MIRE host shares one worker and one disk, so the busiest trap could starve all the others: a denial of service we would have caused ourselves.
| Option | Effect on them | Effect on us | Verdict |
|---|---|---|---|
| Block 47.79/16 and 47.82/16 at Cloudflare | They move to other IPs and come back | Quiet logs, nothing learned | Rejected: fixes the symptom only |
| Remove the links | They leave at once | Tree gone, bait gone | Rejected: wastes the catch |
| Slow them down and end the maze | Each step costs up to 90 s, and the trail runs out | Few requests, bounded logs | Chosen |
The last option keeps what makes a honeypot useful: the attacker keeps paying, and we keep watching.
The fix

On 9 October at 15:31 we deployed a change that makes the maze slow and gives it a floor. Shallow pages behave as before. Deep pages hold the connection open, longer the deeper they are.
- Depth tarpit. From 4 path segments on, each request waits 30–40 s, plus 10 s per extra segment, capped at 90 s. The cap keeps us under Cloudflare’s 100 s timeout, so the crawler gets a slow page rather than an error it could learn from.
- A floor. From depth 8 the page is still served, bait and all, but without the three links. The maze is now finite.
- A safety valve. Each hold is cheap, but if 3’000 are already waiting, new requests get the normal short delay. A crawler that fans out wider can never fill The MIRE’s connections and freeze every other host.
- Evidence. Each log line now records depth= and held=, so we can watch the trap work.
What happened next
The crawler went quiet at 14:08 on 9 October, 83 minutes before the fix went live. It came back at 19:02, straight into the tarpit, carrying a queue of URLs 11 and 12 levels deep. And it turns out it only waits 30 seconds for an answer.
# time client event 09:41:45 47.79.15.54 GET /test/status.php/health/…/status.php depth=12 09:42:14 47.79.15.54 hung up after 29 s status=0 09:43:15 47.79.15.54 page ready after 90 s, nobody left to send it to
Every deep request is now held for at least 60 seconds, so the crawler hangs up every time and leaves with nothing: no page, no bait, no new doors. Its daily count fell from 41’233 on 9 October to 8’947 on 10 October, and 2’030 by Sunday morning. With no links left to find, its queue can only shrink. On our side the cost is tiny: at the busiest moment about 40 requests were waiting at once, against a safety valve of 3’000.
Lessons for anyone building a honeypot
- Design for the visitor who never stops. A link built from request.path multiplies with every click. No one follows it forever, until something does. Give every maze a floor.
- Rate-limit by behaviour, not by IP. Thousands of rotating addresses defeat per-IP counters. Path depth gave this crawler away.
- Tarpits must fit the pipe. Hold for less than your CDN’s timeout, and cap how many holds you park, or the trap becomes your own outage.
- A trap that costs you more than the attacker is not a trap. Measure your own logs and disk as closely as theirs.



A crawler on Alibaba Cloud in Singapore spent three weeks walking deeper and deeper into a fake “staging server” on cfd.mire.cc. By 9 October it was making more than 41’000 requests a day, doubling about every two days, all into a maze that had no exit. This is the story of how it got in, why nothing stopped it, and how we kept it stuck without letting it run on forever.
The MIRE is a honeypot: it serves convincing bait to bots that probe for secrets. In What 421’000 dead-ends taught us we started building traps around what attackers actually ask for. Here the trap worked almost too well. One crawler followed the maze tirelessly, and the honeypot’s own success became the risk.