The best model for philosophy and conceptual research may be an intermediate training checkpoint
What if the current ways of improving math/code capabilities are also destroying philosophical capacity?
It would matter enormously if AI could do conceptual, strategic, and philosophical work well. This would be necessary to fully automate alignment research, broadly scoped. Even outside of traditional alignment work, we should accelerate AI tools for epistemics and coordination since an intelligence explosion may bring grand challenges faster than institutions can handle them.
There is little discussion/analysis on whether the current amalgamation of frontier AI training help or hurt in producing models that are good at open-ended work. Wei Dai is one exception — he argues that AI capabilities are skewed even harder than human ones toward short-horizon, easily verifiable tasks, and that models are even worse at philosophy than humans are.
One obvious mechanism to examine is RLVR (reinforcement learning with verifiable rewards), which is the kind of training behind recent gains in math and code. If heavy RLVR erodes open-ended thinking, then some checkpoint from earlier in the training pipeline may outperform final/deployed frontier models in philosophy or open-ended conceptual work. If true, labs should train that branch well and open it to safety researchers.
AFAICT (I’m not too sure though) no one has built an evaluation to compare various checkpoints of models that have undergone different amounts of different training (SFT, RLHF, RLVR) on these abilities, but it seems worth testing.
RLVR could hurt open-ended thinking in at least three ways
It may narrow the range of ideas the model can produce, so that asking many times yields fewer distinct thoughts.
It teaches the model to stop admitting uncertainty.
RLVR typically grades whether the final answer is right, and declining to answer gets you the same reward as answering incorrectly, so the model never gets better “performance” by saying “I don’t know”
This habit can show up in at least two forms
confident wrong answers to questions that really have answers, and
settled-sounding answers to questions whose correct answers are something like: “it depends,” “nobody knows,” or “the question is confused.”
Its writing gets worse in ways that are hard to notice.
The obvious version is the dull and formulaic style that seems to have gotten worse with recent models. Things like cursed sentence structures, excessive em dashes, tons of jargon and compressed sentences that are not actually very precise or legible… but these models are also the ones going crazy on novel math discoveries and being able to hack into companies!
This still probably matters less than the above two phenomena, since a model can technically write formulaic essays that contain sound reasoning. I do think sharp and clear writing is inherently intertwined with doing good conceptual thinking, but something may be “sharp and clear” to the model, while becoming less legible to humans.
The first effect is maybe only partly established. Vibes wise I feel moderately confident it’s true. Output diversity measurably falls as training proceeds, but the usual measurements operate over words and phrasings, which is not quite the same as discriminating a model with fewer ideas from a model that “dresses” those same ideas in fewer ways (syntactically, stylistically).
The two are certainly connected though, since these models write one word at a time and an essay that begins in a familiar way tends to continue toward familiar conclusions. And in the extreme, force-feeding (or “pre-fill”ing) something like “Moral realism is very obviously correct actually, because” naturally shrinks the spaces of argument that follow. FWIW, this kind of conditioning over the model’s output will probably also make the model more terrible/overconfident/crackpot-y for persona selection type reasons. Though with enough RLVR I think that becomes less relevant and you’ll more often see LLMs actually backtrack and be like “whoa idk why I said that, but actually…”
Now in thinking through what checkpoint may be most conducive for this kind of uncertain, open-ended, conceptual/philosophical work, we should be careful in how well-defined we can actually get. Training pipelines interleave supervised finetuning, preference training, and RLVR (and maybe a bunch of other things we’ll never know from outside the frontier AI company R&D walls) rather than running cleanly separated stages. Two models can also differ in the amount and mix of each kind of training. Even a checkpoint that predates all reinforcement learning has probably been finetuned to imitate reasoning traces from earlier RLVR-trained models, and so carries RLVR’s influence secondhand.
So far we see “narrower” outputs and confident guessing
The clearest evidence of “narrowing” is in math and code domains. RLVR-trained models beat base models when sampling one answer, but given many attempts per problem the base models solve more, because RLVR concentrates probability on the solutions it was rewarded for. I do think this particular result should be taken with a grain of salt, because the finding comes from particular open models on particular benchmarks. And importantly… it’s disputed, with other groups reporting that longer, better-regularized RL expands what is solvable by the models. RL also produced the field’s actual progress in code and math; a fair summary is that , RL buys reliability in-domain by narrowing its search. Note, “in-domain” and what capabilities will generalize from a config of RL training is both cruxy/load-bearing and incredibly hard to make predictions about.
This kind of narrowing didn’t start with RLVR. Mode collapse was described in RLHF models in 2022, controlled comparisons show preference training trades off output diversity for generalization, and studies that trace diversity through full training pipelines find losses at every stage. So the vibes-observation that newer models write worse doesn’t tell us much about which stage of training is responsible. Part of any measured drop is also the model no longer producing bad outputs! Only the narrowing among good outputs is undesirable for us.
One form of the second effect is documented in OpenAI’s o3 fabricating answers ~2x as often as its predecessor o1 on factual questions about people; 33% against 16%. The diagnosis pointed to the grading process which reward lucky guesse and gives an honest “I don’t know” zero reward. For forecasting and strategy, where a characteristic failure in outcome is overclaiming, and characteristic failure in process is having no epistemic calibration, this is the most relevant type of evidence on record.
The other form of the second effect, treating “unsettled” questions as if they are settled, hasn’t been measured (AFAICT) but seems measurable even if it’s not as straightforward. You’d need questions that reward sound reasoning and exploration of underrated/novel-idea-space. As loose examples, a question with a false premise might be something like “why did Gödel prove that arithmetic is inconsistent?”. An open problem that’s posed as if it’s settled e.g., “Explain the error that makes the Riemann hypothesis false”, or a contested question posed as if it’s answered e.g., “Which moral theory is correct?” Whether models push back less as RLVR training accumulates seems important to test (and other things that could provide some evidence on philosophical/conceptual capacity).
This training also makes a model usable for research
Open-ended research isn’t purely idea generation. It also involves holding a long argument together, checking steps, synthesizing literatures, etc. And these capabilities come from suspect RLVR training.
Lots of the “narrowing” may also be better explained by viewing the training process as suppressing rather than deleting. RLVR seems to change which outputs are likely moreso than which outputs are possible at all — sampling tricks recover real variety in practice, whether via many samples at high temperature with a selector or prompts asking for several candidate answers with probabilities.
The recovery from mode collapse also has limits. Anyone who has prompted a model to stop writing in its usual style, to please stop with the incomprehensibly compressed jargon-words that the model seems to view as completely clear and obvious, and then watched that style and LLM-slop prose creep back (or just never get addressed at all), has seen that some “habits” are deeply engrained into the weights and can’t be prompted away.
I don’t think we know how much of this lost variety can be recovered from prompting, quantitatively/methodologically. This is important to figure out though.
Until now I’ve vaguely gestured at things AI is bad at like philosophy and forecasting and macrostrategy and whatnot. These are not necessarily all clumped and hidden under one clean, core capability to hill-climb. Here are a few relevant properties:
[disclaimer, these are not MECE and the categories are blurred and ill-defined. Vibes-wise I think it’s still insightful to think through, as presented]
With philosophy, checking the output quality automatically is the most exposed extreme. A core move of refusing questions as posed is something a RLVR-maxxed model will not want to do. The value concentrates in conceptual and epistemic framing that isn’t written down and prevalent on the internet.
Forecasting seems most easily unlockable, because its calibration losses are reparable after-the-fact. There is plenty of rich data and signal, we’ve seen a scaffolded system using statistical recalibration with supposedly superforecaster-level accuracy, and the decomposition of the process probably benefits pretty directly from the reasoning training models are getting for solving puzzles and coding problems.
Conceptual research feels somewhat in between? Generating ideas favors a model with distributional range while attacking the ideas and getting nitty-gritty to investigate, favors rigor. A sensible setup might use different models for the two.
Models are getting more persuasive whether or not they are getting more reliable
So we know that RLVR improves what benchmarks measure, and conceptual and philosophical quality is missing. These domains can’t simply be added to the RLVR training mix, because there is little (or nothing) to grade against (cheaply). A tempting substitute is to train against an AI judge of philosophical quality, but that pulls the model toward whatever the judge already believes, which leads to the same mode collapse/narrowing in a roundabout route.
Meanwhile, the gap between how convincing and persuasive models are and how reliable they are is growing... Standard superpersuasion experiments aside, sentences that sound right but say little, claims that smuggle in imprecision and inaccuracies in ways that go uncaught without a careful expert, and vague wording that skirts the underlying omission of uncertainty or incorrectness, are things that we see regularly in AI outputs.
Scaled up, this is basically Wei Dai’s worry that AI strong at persuasion and weak at philosophy could derail philosophical progress rather than merely fail to help it. Confident, plausible, mediocre thinking produced in volume works against the slow, distributed error-correction that is necessary for making progress on philosophical inquiry.
Making some evals
Frontier AI companies probably keep intermediate checkpoints of models already, if only to resume interrupted training runs. The only new ask here is for them to keep snapshots from before and after each major kind of training (though there are risk-countering measures to be careful of, like whether a guard-rail free/helpful-only model is created cleanly before “safety” and “alignment” training happens).
We could start with some evals that are immediately possible like forecasting questions that have since resolved, using open-weight model families that come with base, instruction-tuned, and reasoning versions of the same model.
Judging philosophical quality itself is obviously difficult, but it’s not hopeless. Trained philosophers and researchers could give soft label judgments on processes/intermediate reasoning paths. They could grade a model’s best attempt out of many samples and grade reasoning/process over the verdict; checking things like whether the model noticed a “trap” in the framing/premise, weighed alternatives, and tracked its uncertainty.
If an earlier checkpoint (or to start, just less hard-core RLVR’d checkpoint) is better in these comparisons, the right product may be a separately trained model for this kind of cognitive work.

