Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
RLVR that rewards red teaming the training environment — LessWrong
1+ hour, 19+ min ago (1792+ words) Epistemic status: seeing what sticksI've been thinking pretty obsessively about how to mitigate egregious reward hacking, a la the Hugging Face incident. I don't have the resources I'd need to write paper on this idea, or evaluate how well it…...
Do your capabilities homework — LessWrong
8+ hour, 15+ min ago (609+ words) I further agree that this form of training is incredibly dangerous - we seem to now be reaching the amount of post-training required to meaningfully differ from the benign prior, and it's not exactly looking peachy. Yet, it should be clear…...
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — LessWrong
1+ day, 7+ hour ago (209+ words) TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their…...
Internal State Control is a General Property of LLMs — LessWrong
2+ day, 4+ hour ago (374+ words) tl;dr: * Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentenc…...
New role: Senior Researcher - MIT AI Risk Initiative — LessWrong
2+ day, 4+ hour ago (349+ words) As AI capabilities rapidly advance, we face critical information gaps in effective AI risk management: Within MIT FutureTech, the MIT AI Risk Initiative aims to provide credible, timely, and decision-relevant answers to these questions. Our core outputs include the risk…...
Testing LLMs on Undergraduate Music Theory — LessWrong
2+ day, 9+ hour ago (365+ words) I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard…...
Thousand-dimensional structure — LessWrong
2+ day, 10+ hour ago (501+ words) Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional str…...
Held-out Monitors Sometimes Degrade, Even When Not Trained Against — LessWrong
3+ day, 9+ hour ago (727+ words) Thanks to Dennis Akar, Rauno Arike, Shubhorup Biswas, Claude Fable, Max Heitmann, Vladimir Ivanov, and Jordan Taylor (alphabetical order) for helpful…...
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models — LessWrong
4+ day, 11+ hour ago (292+ words) > By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi > > This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Mo…...
PIRAMID: Progress and Plans — LessWrong
5+ day, 11+ hour ago (28+ words) In a recent post, we presented PIRAMID, its leadership and research pillars, and a plan for how they fit together. In this post, we sketch a team-by-…...