10-day longest streak
Rafael (Ralph) Peña Marine vet. CEO and co-founder of WeCheck AI (AI-driven due diligence). Miami. I build with AI agents and I don't trust claims without evidence. Including my own.…
Rafael (Ralph) Peña
Marine vet. CEO and co-founder of WeCheck AI (AI-driven due diligence). Miami.
I build with AI agents and I don't trust claims without evidence. Including my own.
The receipts system:
- rules-with-receipts — a Claude Code quality pack that ships with its own eval evidence, including the tests where it does nothing
- rulebench — point it at any rules file, get honest behavior deltas (its first published run declined to confirm our own headline, and its six-pack study reported one of its own rows as grader noise, with proof; that's the point)
- agent-failure-modes — the AFM Index: numbered failure modes of AI coding agents, with detection traps and evidence grades
Start here:
- About to let an agent into an unknown repo? →
pipx install agent-zero-trust, thenazt scan .(agent-zero-trust). - Want the operating discipline? → Install rules-with-receipts (60-second bootstrap in the README).
- Want to test your own rules file? →
pipx install rulebench(rulebench). - Want the failure taxonomy? → Read the AFM Index (start with "the five you'll see this week").
Find me: @Ralfyishere · LinkedIn
-
rules-with-receipts ★ PINNED
A quality pack for Claude Code that ships with its own eval harness and honest A/B evidence — rules with receipts, not vibes.
Shell ★ 2 20d agoExplain → -
agent-zero-trust
Zero-trust repo intake for AI coding agents — scan the instruction environment before Claude Code, Cursor, Codex, or Gemini touches a repo. Ships its own false-negative ledger.
Python ★ 4 21d agoExplain → -
piensalo
PIÉNSALO is the open artificial cortex for AI—reduce context tokens, verify response quality, expand when needed, and safely fall back.
Python ★ 1 13d agoExplain → -
agent-failure-modes
The Agent Failure Modes Index (AFM): a numbered taxonomy of how AI coding agents fail — with transcript signatures, detection traps, interventions, and evidence grades.
★ 0 21d agoExplain → -
rulebench
Does your CLAUDE.md actually do anything? Trap-test your agent rules in isolated sessions and get honest deltas.
Python ★ 0 21d agoExplain → -
ralfyishere
The receipts stack for AI coding agents: intake scanning, tested discipline, honest evals, failure taxonomy.
★ 0 23d agoExplain →
No repos match these filters.