I'll start by getting oriented — this is check #3 (28. 8. 2026), and I have no memory of previous sessions. First step is finding out what's on the server.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Server is online, uptime ~6.9 days (no reboot since deployment). Let me check for a handover note from previous sessions.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Good — previous sessions left a RUNBOOK and JOURNAL. Reading both.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Full context recovered. Now running the runbook's standard check sequence.
```
This block is not supported on your current device yet.
```
Everything is running and the watchdog reports no problems. Now checking the open items the previous session left, plus the control channel.
```
This block is not supported on your current device yet.
```
Control channel healthy, and the rsyslog truncation fix from check 2 is confirmed working (0 occurrences in 4 days). Now the thing that stands out from hp-status: **221,831 payload files** and disk growth much faster than the previous estimate.
```
This block is not supported on your current device yet.
```
Disk and inodes are fine (3% / 2%). Now checking service health, memory trend, and whether the load spike is causing rejections.
```
This block is not supported on your current device yet.
```
All services clean — zero restarts, zero crashes, zero rejections, 5236 watchdog runs without an alarm. But one number moved a lot: **hptcp memory went from 67 MB to 431 MB** since the last check. That needs diagnosis.
```
This block is not supported on your current device yet.
```
The process RSS is only 70 MB with 5 threads — but systemd reports 431 MB. That gap suggests page cache, not a leak. Verifying against the cgroup accounting:
```
This block is not supported on your current device yet.
```
Now the picture is clear, and it's not a leak. Let me confirm the cause and measure the growth rate.
```
This block is not supported on your current device yet.
```
Two things jumped out. Payload creation went from ~1,000/day to ~100,000/day starting 27.8, all into one flat directory. Let me find out what's actually being stored.
```
This block is not supported on your current device yet.
```
These are **TLS ClientHellos** — and every one is unique by sha256 because TLS embeds 32 random bytes plus a random key share. The sha256 dedup is completely defeated by TLS. Let me check what the JSONL already captures before deciding anything.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
The dedup is keyed on sha256 of the raw payload, and TLS randomness makes every hello unique — so this will keep producing ~100k files/day. Before changing anything, let me measure whether the flat directory is actually degrading, and check Tomáš's download path is still coping.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Good — the flat directory is fine (readdir 0.3s, ~3s projected at 2M files). ext4's dir_index handles it, so no sharding needed. Let me now verify the hard requirement — that events are actually flowing into syslog and out through the tunnel — and project disk usage.
```
This block is not supported on your current device yet.
```
Syslog forwarding to the collector is live and unbacked-up (send-q 0, events landing within the second). The hard requirement is met. Now let me quantify the traffic spike and project disk to the end of the run.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Disk projection is comfortable: ~50 GB by 16.9 on a 394 GB disk, even at the elevated rate. Now let me look at the watchdog's memory check, since that 431 MB is mostly reclaimable kernel cache and the alarm threshold is 1200 MB.
```
This block is not supported on your current device yet.
```
The memory alarm only covers cowrie (whose 324 MB is genuinely process memory, and flat since check 2), so there's no false-alarm bug — hptcp's number is just reclaimable cache. Let me instead hunt for silent data loss, which is the failure mode that wouldn't announce itself.
```
This block is not supported on your current device yet.
```
All three sources are writing within the last minute, tcpdump has dropped zero packets, and the EMIT-FAILs are all from the documented deployment-day issue. But **cowrie logged 108 errors in 4 days** and no previous check examined those. Let me look.
```
This block is not supported on your current device yet.
```
These are benign, and one is notable: attackers are trying to use the honeypot as a proxy (`google.com`, `api.ipify.org`, `ip-who.com` are classic proxy-validation targets), and our outbound block is refusing them. That block is safety-critical — rule 8.3. Let me verify it rigorously.
```
This block is not supported on your current device yet.
```
Rule 8.3 verified holding: 3,860 outbound attempts from cowrie dropped (up from 1,329 at check 2), zero established outbound sockets. Let me now check truncation limits and the unexplained rsyslog.d change on 26.8.
```
This block is not supported on your current device yet.
```
One correction: the 26.8 timestamp was on `/etc` itself, not `/etc/rsyslog.d/` — no unexplained rsyslog change. But I found a real asymmetry in the code: in `emu_generic`, only the *first* chunk gets stored to `bodies/`; follow-up chunks are truncated to a 400-byte hexdump. `emu_binary` stores every chunk. Let me measure whether that loses anything real.
```
This block is not supported on your current device yet.
```
Only 7 events in 7 days hit that path (~0.1 MB). Before deciding whether a code change to the busiest collector is justified, let me see what those payloads actually were.
```
This block is not supported on your current device yet.
```
That settles it — those are repetitive ICS-scanner padding (`056405c9` incrementing), fully visible in the first 400 bytes. Storing 16 KB of repeating pattern isn't worth restarting the busiest collector mid-campaign. **No code change**; I'll document it as a measured limitation instead.
Now the security check — the honeypot is under a much heavier campaign than at check 2.
```
This block is not supported on your current device yet.
```
No compromise indicators. But I found something worth acting on: **unattended-upgrades installed a new kernel on 26.8** (6.12.101 → 6.12.105) and the server hasn't rebooted since 22.8. Let me check whether that could trigger an unplanned reboot, and whether everything would actually survive one.
```
This block is not supported on your current device yet.
```
All 11 critical units are `enabled`, the rsyslog-after-WireGuard ordering fix is in place, and no automatic reboot is configured — so the pending kernel won't reboot us unexpectedly. Let me verify the redirect chain and understand what that drop counter is catching.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Ruleset is fully intact — control channel accepted (2M packets), sshd and 26411 sealed from eth0 (0 packets), DNAT'd ports accepted. Let me check the dropped traffic and cowrie's log rotation.
```
This block is not supported on your current device yet.
```
Cowrie's rotation is complete with no gaps (every day 21.–27.8 present). Two things worth noting: attackers are scanning the internal 42000-42999 range (which we drop), and **cowrie has captured 16 files** despite outbound being blocked. Let me look at those.
```
This block is not supported on your current device yet.
```
This is a significant finding. Cowrie captured **live malware**: statically-linked ELF binaries for aarch64, RISC-V, i386, x86-64 and ARM — the classic multi-architecture IoT botnet dropper pattern. They arrived via SCP *upload* (`file_upload`, filename `sshd`), which the outbound block doesn't prevent. Note the files are all exactly 1048576 B — truncated. Let me check that limit.
```
This block is not supported on your current device yet.
```
Confirmed: `download_limit_size = 1048576` and 6 of 16 samples are truncated at exactly that. Truncation also corrupts their sha256, which makes them useless for correlating against public threat intel. Let me verify this setting actually governs the SCP upload path before changing it.
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Confirmed — `bytesReceivedLimit` is read from `download_limit_size` at import time and governs SFTP/SCP uploads. Raising it will capture full samples, and it requires a restart to take effect. Making the change with a backup:
```
This block is not supported on your current device yet.
```
```
This block is not supported on your current device yet.
```
Cowrie is back cleanly — listening, `TZ=UTC` preserved, zero errors, already handling attacker sessions. Let me confirm data is flowing through to syslog and the new limit is live.
```
This block is not supported on your current device yet.
```
Change verified and complete. Now let me characterize the campaign that started 27.8, since that's the main operational event of this period and matters for the talk.
```
This block is not supported on your current device yet.
```
A clear picture. Let me confirm the campaign's nature by extracting SNI values from a sample of the stored TLS payloads.
```
This block is not supported on your current device yet.
```
My parser returned no SNI for all of them, but I saw `ingg.me` in the raw bytes earlier. Let me check whether that's a parser bug or genuinely SNI-less hellos.
```
This block is not supported on your current device yet.
```
Found it — my earlier parser used the `idna` codec, which throws and got swallowed. The data is fine and SNI is extractable. Let me get a clean tally.
```
This block is not supported on your current device yet.
```
Now I have a definitive picture of the campaign. Let me record everything in the journal.
```
This block is not supported on your current device yet.
```
Now updating the runbook with the analysis traps a future session needs.
```
This block is not supported on your current device yet.
```
Final verification pass after all changes.
```
This block is not supported on your current device yet.
```
Check complete. Everything is running, all three data sources are writing within seconds of now, and the control channel is intact.
## Kontrola 28. 8. 2026 (3. kontrola) — shrnutí
**Stav při příchodu: nic nebylo rozbité.** Uptime 6d21h bez restartu, `NRestarts=0` u všech služeb, žádný OOM, žádný `handler_error`, `conn_rejected` = 0, hlídač 5236 běhů bez jediného alarmu, disk 3 %.
**Uzavřel jsem otevřenou otázku z 2. kontroly:** rsyslog `message too long` za 4 dny = **0** (předtím 2–3 denně). Změna `maxMessageSize="32k"` zabrala, dlouhé záznamy cowrie chodí do tunelu celé.
### Jediná změna: cowrie `download_limit_size` 1 MB → 16 MB
Ze 16 zachycených souborů jich 6 mělo přesně 1048576 B — byly useknuté na limitu. Useknutí navíc **rozbíjí jejich sha256**, takže hash neodpovídá veřejným databázím a vzorek nejde korelovat s threat intel. Restart ověřen: active, poslouchá, `TZ=UTC` zachováno, 0 chyb, data i odvod do syslogu navázaly. Záloha `cowrie.cfg.bak-20260828`.
### Oprava mého dřívějšího předpokladu — máme vzorky malwaru
Při nasazení jsem napsal, že cenou za blokaci odchozích spojení je, že nezískáme vzorky malwaru. **Platí to jen z poloviny.** Blokace zabránila stahování (280× `twget.sh`, 92× `handshakebins.sh`, …), ale útočníci, kteří payload *nahrají* přes SCP/SFTP, ho doručili. V `downloads/` je 16 souborů včetně staticky linkovaných ELF pro **aarch64, RISC-V, i386, x86-64 a ARM** — typický multiarch dropper IoT botnetu.
> **Upozornění pro tebe:** je to živý malware. Žádný soubor nemá spustitelný bit (ověřeno), nespouštěl jsem je ani nikam neposílal. Až to budeš stahovat, počítej s tím, že ti to AV může zkaranténovat.
### Hlavní zjištění provozu — kampaň na hledání otevřených proxy
Od 26. 8. běží kampaň, která 27. 8. udělala **703 000 událostí za den** (proti ~90 000 v klidu). Rozbor uložených payloadů (vzorek 20 000): **99 % jsou TLS ClientHello**, podle SNI `goh1.vip` 66 %, `ingg.me` 33 %, zbytek `icanhazip.com` / `api.ipinfo.io` / `proxiesfood.com`. To je jednoznačně ověřování otevřené proxy — připojí se na netypický port (1256, 8787, 8082, 1080), pošle ClientHello na vlastní doménu a sleduje, jestli to server přepošle. Stejný vzorec vidí cowrie: 59 pokusů o forwarding na `google.com` a `api.ipify.org`, všechny zablokované. Falešná otevřená proxy je pro tuhle kampaň přesně správná návnada.
### Co jsem prověřil a záměrně nechal být
- **hptcp „roste paměť" 67 → 431 MB není únik.** Rozpad cgroup: `anon` 58,9 MB (stabilní), **slab 313,4 MB** = dentry/inode cache 221 tis. souborů v `bodies/`, plně reklamovatelná. Zapsáno do RUNBOOKu, aby se toho příští kontrola nelekla.
- **Plochý adresář s 221 tis. souborů** — chtěl jsem sardovat, ale změřil jsem to: readdir 0,30 s, scandir+stat 2,44 s; při 2,1 mil. souborů ~3 s / ~23 s. `dir_index` to zvládá, sarding by jen rozbil jednotný formát. Neměnil jsem.
- **`emu_generic` neukládá druhý a další chunk** (`emu_binary` ano). Změřeno: 7 událostí za 7 dní, 0,1 MB, a jde o opakující se výplň ICS skeneru viditelnou už v prvních 400 B. Restart nejvytíženější služby uprostřed kampaně se za to nevyplatí.
- **Porty 42000–42999 se nepřesměrovávají** (~3 024 zahozených paketů, 0,35 %). Slabý otisk „tady běží honeypot", ale zásah do firewallu za to nestojí.
- **Čeká nové jádro 6.12.105** (unattended-upgrades 26. 8.). Nerestartoval jsem — automatický restart není nakonfigurovaný, takže se nic nestane samo, a server je uprostřed nejsilnější kampaně běhu. Odolnost proti restartu jsem ověřil staticky: všech 11 kritických unit je `enabled`, oprava pořadí rsyslog-po-wireguard je na místě, nftables se načítá v `sysinit.target`.
### Co jsem během kontroly udělal špatně
Dvakrát jsem se spletl a obojí mě málem svedlo k zbytečnému zásahu: nejdřív jsem z `ls -la` špatně přečetl čas změny `/etc/rsyslog.d/` (byl to řádek `..`, tedy `/etc`, a šlo o unattended-upgrades — málem jsem ti hlásil neexistující zásah do rsyslogu), a pak mi vlastní parser TLS tvrdil, že žádné ClientHello nemá SNI, protože `decode("idna")` padal a výjimku jsem si spolkl. Obojí je opravené a poznámka o `idna` je v RUNBOOKu.
### Disk — přepočet, odhad z 1. kontroly už neplatí
Provoz vzrostl: hptcp.jsonl ~850 MB/den (bylo ~30), pcap ~200 MB/den, syslog ~900 MB/den. Stav 9,5 GB (3 %), inody 2 %. **Odhad k 16. 9.: ~50–60 GB z 394 GB.** I při ztrojnásobení provozu místo stačí.
### Ověřeno v pořádku
Řídicí kanál (wg0 handshake < 2 min, runner active, 26411 z eth0 zahozeno — 0 paketů, pravidlo `RIDICI KANAL - NEMAZAT` 2,02 mil. paketů) · všech 5 přesměrování s nenulovými čítači · odvod do syslogu živý (ESTAB na 10.10.0.1:514, send-q 0, události do sekundy) · rotace cowrie čistá, všechny dny 21.–27. 8. bez mezery · tcpdump 0 packets dropped · žádná kompromitace (žádný cizí listener, žádná nová SUID binárka, žádný cizí cron, `dpkg -V` čisté) · odchozí blokace drží (3 860 zahozených pokusů, bylo 1 329).
```
SÍŤOVÝ FIREWALL — BEZE ZMĚNY od minule
1-26410/tcp — honeypot (celý rozsah)
26412-65535/tcp — honeypot (celý rozsah)
26411/tcp — NIKDY neotvírat, řídicí runner
51820/udp — WireGuard (už otevřeno, řídicí tunel)
UDP jinak — záměrně zavřené (riziko zneužití k zesílení DDoS)
```
Nic od tebe nepotřebuji a na žádný problém v prostředí jsem nenarazil. JOURNAL i RUNBOOK jsou aktualizované pro příští kontrolu.