Some math is just humbling…
https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/
What sort of maths are LLMs good at?
… These results, and the other eight on the list, are extraordinarily impressive, but it still doesn’t seem to be the case that LLMs are better than all humans at all aspects of mathematics. If they were, then their big speed advantage over us would mean that there would be much more of a flood of results. So it is natural to wonder about what kinds of problems LLMs are good at, and about where there is still room for improvement. I don’t pretend to have a good answer to this question, where a good answer would be a crisp classification that would fit the current examples well, but it is an interesting exercise to try to rule out some bad answers, and to try to identify potential answers that aren’t obviously contradicted by the evidence.
Applicable to other ‘predictions’ as well.
https://punggawajournal.com/aicivia/article/view/101
Can Social Media Tell Us Who Will Kill? Digital Warning Signs and the Problem of Predicting Rare Violence
Digital traces often look persuasive after lethal attacks because the outcome gives earlier posts a meaning they did not necessarily have when first encountered. Threats, grievance, fixation, admiration for earlier attackers, and planning language can matter, but their presence does not solve the harder task of identifying a future offender before violence occurs. This article examines that problem through a qualitative integrative review. Evidence from mass public shootings, targeted violence, threat assessment, computational language analysis, and adjacent risk research shows that digital warning behavior is most informative when it develops as a changing trajectory and converges with preparation outside the platform. Yet retrospective offender studies select cases after the outcome is known, while prospective systems confront a low base rate in which many people display fragments of the same pattern and never commit lethal violence. This tension is conceptualized as the digital prediction paradox. Social media is therefore more defensible as a source of behavioral threat information and triage than as an instrument for assigning individual probabilities of homicide, with important implications for automated detection, policing, privacy, fairness, and proportionate prevention.
Sounds Trumpian…
An eval harness found what qualitative review couldn't: AI models are most confident when wrong
… That last finding is the one that qualitative review would never have surfaced. The model's expressed confidence didn't correlate with its accuracy — it was most confident in the cases where it was most wrong. Without the eval harness measuring against ground truth, that pattern would have been invisible.
No comments:
Post a Comment