Commit graph

3 commits

Author SHA1 Message Date
Hound
2cf9a740c3 rules: every rule needs an anchor, and the test now proves it for all of them
Stage 3 passed on throughput — 27,339 events, zero rescues, Caddy
unmoved — and then the soak found what the load test could not.

Two more false positives, both the same shape as the ones before:

* EICAR matched an 8.5 MB rustc incremental-compilation cache, because
  the test source being compiled contains the literal. Hound moved it to
  quarantine mid-build and rustc panicked. The standard defines the
  EICAR file as exactly that 68-byte string, optionally padded to 128,
  so the rule now says filesize <= 128.

* The webshell rule matched a 4.3 MB AI session transcript, because the
  conversation had been discussing webshells and therefore contained
  "<?php", the eval pattern and "$_POST". The transcript was moved to
  the vault and its history lost. A webshell is a PHP file: small, and
  opening with a PHP tag. Now filesize < 1MB and $php in (0..4096).

The interesting part is why the second one happened at all. After the
first, I added a test asserting that a large file containing rule
strings is not a threat — and hand-listed the strings. I listed the
miner's and the rootkit's and forgot "<?php". The test passed and the
transcript was quarantined anyway.

So the test now extracts every string literal from the rule pack itself
and builds the haystack from those. A rule added tomorrow is covered
without anybody remembering to cover it. It also asserts the extractor
actually found the strings, because a parser that silently returns
nothing would make the whole thing vacuous.

Both fixes have a paired test that the detection still works: a real
68-byte EICAR file is caught, padded to 128 it is caught, and a real
webshell is caught.

Worth recording, because it is not a bug: six houndd tests failed while
the gate was armed. Hound quarantined the EICAR fixtures the test suite
had just written — correct behaviour, colliding with a suite that
creates real malware samples. Running the antivirus's own tests on a
gated machine needs thought; the tests are not wrong and neither is the
gate.

The definitions chain now works end to end: pack built from OSV, signed
with the release key, published to /srv/houndav/defs, installed, and
verified on load against the public half compiled into the agent —
"defs: 19 indicators from 1 pack(s) [2026.08.21]".

The public key is in the source on purpose. The agent is open source and
anybody should be able to check that the definitions they received are
the ones we published.

300 tests pass. Gate is off pending these fixes being soaked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 08:37:49 -05:00
Hound
be5396821d gate: stop blocking reads, and stop calling documents malware
Armed the gate on / on the live server. Aborted after about twenty
seconds. The box was never at risk — Caddy stayed sub-millisecond and
load never rose — but the gate blocked reads of an AI agent's session
transcript, reporting it as Linux.Coinminer.XMRig.

It was not wrong about the bytes. That transcript contains
"stratum+tcp://", "donate-level" and "xmrig" because the miner rule was
being written in that session. The rule matched a document ABOUT
malware.

Three bugs, none of which the tmpfs stage could have shown:

1. The miner rule had no file-type condition, so any text mentioning
   mining tripped it: threat-intelligence reports, security blog posts,
   support tickets, an antivirus's own logs. It now requires ELF magic,
   as the rootkit rule always did. Two regression tests: a transcript
   discussing the rule is clean, and an ELF carrying the same strings
   still matches — the fix must not cost the detection it exists for.

2. The gate requested FAN_OPEN_PERM, so it held every OPEN, not every
   execve. A matching file could not be read by anything. That is a
   different product from the one advertised, and on a multi-tenant box
   it is a denial of service against the operator rather than a defence.

   Read events are no longer requested at all. FAN_OPEN_EXEC_PERM and
   FAN_CLOSE_WRITE cover the threat: execution is refused before it
   happens, and anything malicious written to disk is quarantined when
   the write completes. An interpreted script is caught as it lands
   rather than as it is read — the same protection, one step earlier.
   `serve` also guards deny-on-exec explicitly, so re-requesting read
   events later cannot silently restore the old behaviour.

3. Hound did not exclude its own state. /var/lib/hound and /run/hound
   are now always excluded; the vault holds live malware by definition.

Henry asked whether the single watchdog rescue was queue pressure or
scan time. It was scan time: the gate inherited the on-demand 100 MB
limit and tried to read and match a multi-megabyte transcript inline
while holding a process. A gate's budget is a deadline, not a size, so
it now caps at 32 MB — anything larger is allowed through unread rather
than turned into a rescue, which is a process released unscanned and
worse than never having looked.

Dropping read events made everything faster, because most opens on a
running machine are reads:

  latency     +1.38 -> +0.79 ms per exec
  throughput  2,680 -> 4,178 execs/sec (58% of ungated, was 36%)
  events      1,179 in five seconds on an idle tmpfs -> 1

Re-verified on the tmpfs: an ELF miner is quarantined before it can even
be made executable, a document naming every one of its strings is
readable, and a clean binary runs.

297 tests pass. The gate stays off; stage 3 gets attempted again with
these fixes and fresh numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 08:13:02 -05:00
Hound
6746182f18 houndd: replace the clamscan fork with yara-x in process
The old engine shelled out to clamscan for every scan, and clamscan
reloads a 169 MB signature database on every invocation. Measured on a
68-byte EICAR file: 6.5 seconds and ~1.5 GB RSS — paid once per file,
and realtime.rs called it once per inotify event.

Replaces it with HoundEngine: yara-x compiled once at daemon start,
held in memory, one scanner reused across a whole walk, plus a verdict
cache keyed on (dev, ino, mtime, size) so an unchanged file that has
been seen before never reaches the matcher.

Measured after, same machine, same EICAR file:

  single file      6.5 s  ->  4 ms
  400 files cold      --  ->  9 ms
  400 files warm      --  ->  5 ms

Also here:

- rules.rs: hot-swappable rule store. Built-in pack is embedded so a
  fresh install detects something before it has ever reached the
  network; on-disk packs load from $HOUNDD_RULES_DIR, /var/lib/hound
  or the XDG data dir. Reload swaps an Arc, so in-flight scans are
  never torn out from under.
- cache.rs: bounded FIFO verdict cache. Any of the four key fields
  changing means rescan, so edits, truncates and replace-by-rename all
  correctly miss.
- The goodware gate: every rule is scanned against all of /usr/bin,
  /bin and /usr/sbin in CI, and a single hit fails the build. It has
  already earned its keep — it caught a reverse-shell rule that matched
  /usr/bin/sudo, which is now removed rather than tuned. A rule that
  quarantines sudo is worse than no rule at all.
- ScanEngine is Send + Sync and selection stays per-call, so
  HOUNDD_ENGINE=clamav still reaches the legacy path for comparison.
- ScanResult.skipped reports files passed over for size instead of
  quietly counting them as clean.
- Settings gain theme (auto/light/dark), tray_icon_style (color/mono),
  close_to_tray and confirm_quit, normalised daemon-side because
  clients are not trusted to send a theme we can render.

57 tests pass, up from 29.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:05:09 -05:00