Hacker Newsnew | past | comments | ask | show | jobs | submit | babelfish's commentslogin

They have Astra in other benchmarks lower on the page. They just don't want to show it winning

The chart is cursorbench though and they asked about the "deceptive graph"

this is exactly it.


Are there intentionally zero skills listed/curated?

I have a backlog of 20 or so to review right now and am getting that out as fast as I can.

you can just ask it to explain those things!

Sure, and it will explain it to you from the same embedding space that yielded the solution that LRU cache is unbeatable.

It'll just be telling you what the data it ingested claims. Not what is necessarily true. It's as subject to garbage in, garbage out as anything, and you don't know what it's actually trained on.

Why is this a 'rug pull'?


Every sentence in this comment is incorrect



A human definitely didn't, but one of the benefits of formal verification is that even if the work done to achieve something is slop-y or excessively verbose, solvers like Lean guarantee that the initial proposition (assuming it was written correctly and in this case was definitely reviewed by humans) is definitively True. This is true across other domains of formal verification outside of math as well


guaranteed, up to lean itself having bugs that are exploited by the LLM :shrug:


Do you have proof of this bug or something? Is this just envy against computers now ?


as mentioned elsewhere, there was a bug in the lean kernel exploited by AI to prove a false statement roughly a month ago

https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...


Got it. Thanks. I feel people are using this single story to downplay this feat. There's definitely a chance but I don't see any indication of similar bugs in here or the openai's proofs that were created a month ago as i think these companies might've vetted it enough and the other team who's working on similar lean proof for this also seems to have acknowledged this feat


I also doubt this is leveraging a lean4 kernel bug, but I also do not think that a 13m LoC proof that has not been human reviewed closes the book on our understanding of Fermat's Last Theorem, in part because of the decided possibility of a kernel bug being used somewhere in those 13m lines.


Of course, there's a possibility but it exists everywhere but there's no sign till now that it has. Same with openai's proofs.


How about all of these bugs from last week?

https://leodemoura.github.io/blog/2026-8-24-postmortem-for-t...

...I'm not saying this FLT result is compromised. I suppose things depend on your perspective where we are on the spectrum of "finding more bugs means there are fewer left to discover" vs. "finding more bugs probably means there are still unexplored corners out there".


Sure. But experts seem to be aware of the direction of those solutions so it seems unlikely there could be some hidden bug which disproves it. But it could be possible.


Well considering the proof is pretty much accepted by mathematicians to be correct (I'll be happy with that!), it would be sort of unnecessary to cheat. Maybe if some aspect is really tricky to formalize it could have done something there? If I had to search for it, I would go for parts of the original proof that are "outsourced" to other mathematical works. Imagine one of the agents struggling to download a paper due to a paywall or whatever and just deciding to cheat lol


must be fixed already?


I'm still getting 404's right now.


I'm still seeing 404s


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: