Posts

Show HN: JevBench, a reproducible benchmark for typed decision models https://ift.tt/aJrpRiv

Show HN: JevBench, a reproducible benchmark for typed decision models Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison. Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on. JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting. A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost. Leaderboard right now: #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes: https://ift.tt/UrTtuA0 Two no-signup demos: https://ift.tt/GuSOEt...

Show HN: Training a model to identify AI web content from structure alone https://ift.tt/6NCSTHX

Show HN: Training a model to identify AI web content from structure alone Hey HN! We’re Vincent and Jochen from Sitefire ( https://sitefire.ai ). We have been working together for years, with backgrounds in RL/optimization at Stanford and software engineering from Technical University Munich (TUM). With Sitefire (YC W26), we help marketing teams get recommended by AI Search (ChatGPT, Google AI Overviews, AI Mode, Claude, etc.). Our software monitors prompts, sees which web pages get cited, and uses these insights to help marketing teams take action, e.g. create YouTube videos or write the right blog posts. This means we have a commercial stake in AI-generated web content. And for now, high-information, AI-generated content works great to get cited and recommended in AI Search. But after talking to hundreds of marketing teams, it became clear that everyone despises AI-generated content (“AI slop”). And yet, everyone still wants to leverage AI to create content. So we asked ourselves: wh...

Show HN: Notes on Agentic AI – A text-first guide for practicing engineers https://ift.tt/jchOJbA

Show HN: Notes on Agentic AI – A text-first guide for practicing engineers https://ift.tt/mF1cJTu September 22, 2026 at 11:49PM

Show HN: Viaduct – C4 models that coding agents can read and update https://ift.tt/7ITcjgZ

Show HN: Viaduct – C4 models that coding agents can read and update https://ift.tt/68h2gLC September 21, 2026 at 11:20PM

Show HN: AURA – Open-source behavioral threat detection for LLMs https://ift.tt/kzJ9FfD

Show HN: AURA – Open-source behavioral threat detection for LLMs https://ift.tt/qVNDYS1 September 21, 2026 at 10:35PM

Show HN: MidiSlayer, a desktop sight-reading trainer for MIDI keyboards https://ift.tt/bMkUynj

Show HN: MidiSlayer, a desktop sight-reading trainer for MIDI keyboards https://midislayer.com/ September 20, 2026 at 10:48PM

Show HN: Jev filters large HN discussions https://ift.tt/XoGpcyf

Show HN: Jev filters large HN discussions https://ift.tt/1fGYKBH September 20, 2026 at 10:37PM