logoalt Hacker News

Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone

81 pointsby NickySlickstoday at 11:06 AM23 commentsview on HN

Comments

voodooEntitytoday at 12:35 PM

The following rant is not against the owner/project - but...

What an irony. I cant publish a attack surface mapping / pentesting tool i wrote which runs fully deterministic and really controlable due to "dual use" legal problems - but llm driven tools hit public space......

sorry for the rant....

show 3 replies
haeseongtoday at 2:09 PM

What does the 50% look like when it fails? Garbage the parser throws out is easy to handle, but a well formed command aimed at the wrong host gets past the scope check, and you would only catch that reading the report afterward.

aaa_aaatoday at 2:44 PM

Judas Priest reference?

kreidematoday at 12:08 PM

I completely forgot that AI can very much also attack networks/devices in the wild. Interesting project.

oquidavetoday at 12:25 PM

Why phone? This cuts out a lot of phones. Why not on a computer?

show 2 replies
imranshah10140today at 3:05 PM

Will this work on iphones as well.

NickySlickstoday at 11:09 AM

I built Nightcrawler, an open-source autonomous penetration-testing agent that runs entirely on an Android phone.

The project started with a question: how much of a real pentesting workflow could I run locally on relatively old mobile hardware, without relying on a cloud model or API?

Nightcrawler runs a 1.2B-parameter model locally on the Adreno GPU of a OnePlus 8. The model chooses targets and tools, while a separate scope-enforcement proxy validates every command before execution. The system maintains per-host memory in SQLite, rotates between targets, matches detected versions against a local CVE database, executes multi-step playbooks, and generates a structured report.

A few implementation details that may be interesting:

Local inference runs at roughly 115 prompt tokens/sec and 13 generated tokens/sec. The small model only produces a usable command around 50% of the time, so much of the engineering is recovery logic, duplicate detection, persistent memory, and deterministic playbooks. Every command passes through a separate scope and safety layer rather than trusting the model to remain in scope. The project includes a dry-run mode, so the agent loop can be tested without executing real network commands or owning the phone hardware. I've had it running on my home network for the past 3 months uninterrupted

show 1 reply
michaelksalemetoday at 3:57 PM

[flagged]