ToolsThe story, in brief

Show HN: We post-trained a model that pen tests instead of refusing

Not a pilot. A YC startup just shipped an AI pen-test tool that actually breaks into systems—something OpenAI and Anthropic won't let their models do.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A post-trained model fine-tuned on CTF data is enabling mid-market companies to run AI-powered security audits without enterprise gatekeeping. The counterintuitive move: removing guardrails responsibly to democratize vulnerability detection.

The key facts

8 to know
  1. Post-trained on decade of capture-the-flag contests using Kimi K2.6 base model

  2. Two modes: read-only security scan + active adversarial pen-test (gated)

  3. Security scan found integer overflow vulnerability in Bank of Anthos demo

  4. Multi-agent swarm architecture with parallel subagents

  5. Free tier: up to 2M tokens; pay-per-token beyond

  6. Local CLI binary; inference API over TLS; free installation

  7. Company: Cosine (YC W23)

  8. SFT on CTF writeups + RL with verifiable exploit rewards

Go to the source

Hacker Newsargusred.com

Publisher excerpt: Anthropic and OpenAI's publicly available models are explicitly guard-railed so that they refuse offensive tasks. And their cyber-focussed models are gated for enterprises. This leaves SMEs and mid market open to major vulnerabilities. AI can be used as both an adversarial and defensive tool in the…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools