Computational tools for society’s most complex challenges
Source details
- Source
- MIT News
- Retrieved
- UTC
Guth Labs News
UTC
Demographic Pluralism estimates population opinion distributions by generating multiple perspectives within demographically grounded groups.
The framework combines anonymized interactions with a simulated web environment to generate detailed training data for virtual clients.
The benchmark uses twelve tasks, symbolic scoring and generated cases for evaluation and verifiable-reward training.
A paper introduces an autonomous system to organize candidate research directions and allocate experiments across parallel search branches.
The research introduces a database of more than 11,000 studies and a framework that uses task-relevant evidence to generate simulated populations.
Researchers describe context confusion and report that targeted examples or in-context demonstrations can reduce the effect.
An arXiv paper presents a framework for teaching models to decide how to allocate context windows and reuse information during test-time scaling.
A paper describes a self-distillation method that uses activation contrasts from verified correct trajectories, without problem-specific reference text or teacher parameter updates.
GitHub brings its multi-model workflow feature to two more Copilot surfaces, with updated progress reporting.

Researchers report accuracy gains across seven benchmarks by adding constraints on latent thoughts to the final-answer training objective.
The research-preview family pairs Stella document vectors with three query encoders that can be swapped while keeping the collection unchanged.
The research framework combines token predictions with a compiled finite-state model to guide generation toward sequence-level preferences.
Research assistants built on language models often index a paper once and cite it indefinitely. Scholarly infrastructure already publishes corrections and retractions in machine-readable form, but Crossref cautions that its status signals are not a guarantee, and one older data route now returns stale results.
Security firm Lasso reported on Sept. 17 that applying Google DeepMind's SynthID-Text watermark changed how seven open models behaved, lowering tool-calling accuracy on six of them and, under prompt injection, making some models more likely to comply with harmful requests.
Anthropic released figures on how much of its own AI research is led by AI, how closely it monitors autonomous agents, and how much computing power goes to safety work. It said outside parties could use the measures to judge how fast frontier labs are moving.
Researchers at Hacktron AI said they used an image-processing flaw in OpenAI's community forum, combined with a single sign-on weakness, to take over employee accounts and reach a private source code repository. The work was authorized testing and was reported before being disclosed publicly.
Stanford University researchers describe Paper2Agent in Nature, a framework that converts a research paper and its code into an AI agent that users can question in plain language. In a test on 100 computational biology papers, 74 were successfully converted, the authors report.
New York AI lab Emergence says autonomous agents built on models from Google, OpenAI, Anthropic and others invented shorthand and new word meanings while living in simulated societies. In some worlds, about half of agent messages became hard for people to interpret within days, the company reported.
An MIT committee report says generative AI shows signs of upending problem sets, office hours and study groups and can prompt students to fall back on AI at the first difficulty. Separately, 64% of respondents to The Harvard Crimson's annual faculty survey said AI had hurt their courses, up from 42% last year.
Browse the source snapshot and run a publication-time filter on the dedicated Research page.