All Articles
16990 articles total
Learning When to Trust in Contextual Social Bandits
The ESA algorithm is introduced to combat 'Contextual Sycophancy' in social bandits by learning per-evaluator trust boundaries through a sparse stream of ground-truth audits.
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
NeSy-Route is a large-scale neuro-symbolic benchmark for constrained route planning in remote sensing, featuring an automated data-generation framework.
ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
ReLope introduces KL-regularized LoRA probes to improve the routing of multimodal LLMs by enhancing the separability of correctness signals in hidden states.
Quantification of Credal Uncertainty: A Distance-Based Approach
A new distance-based approach using Integral Probability Metrics is proposed to quantify total, aleatoric, and epistemic uncertainty in credal sets for multiclass classification.
Mistake gating leads to energy and memory efficient continual learning
Memorized mistake-gated learning is proposed as a biologically inspired plasticity rule that reduces synaptic updates by 50-80% to improve energy and memory efficiency in continual learning.
Action-Aware Generative Sequence Modeling for Short Video Recommendation
The A2Gen modeling paradigm improves short video recommendations by analyzing the temporal dimension of user actions to better capture nuanced preferences.
I Left Google DeepMind
A discussion on Hacker News regarding an individual leaving Google DeepMind.
Microsoft patches bug in video game Age of Empires II
Microsoft has patched a critical vulnerability in Age of Empires II that could have allowed remote code execution via game invites.
FCC to repeal 39% TV ownership cap in boost for Trump-friendly news orgs
The FCC is moving to repeal the TV ownership cap, potentially benefiting specific news organizations.
In memoriam: 7 of our favorite Sam Neill films
A retrospective on the films of actor Sam Neill following his passing.
Third-party app stores coming to Google Play next week as Epic settlement withdrawn
Google Play is introducing third-party app stores following the withdrawal of a settlement with Epic Games.
OpenAI's first branded hardware is... a light-up keyboard?
OpenAI introduces the Codex Micro, a light-up keyboard designed for monitoring agentic threads.
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Research exploring reward models for test-time scaling in LLMs, finding that generative outcome reward models (gORM) are the most robust across domains.
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
Introduction of CrochetBench, a benchmark for evaluating the ability of vision-language models to generate executable crochet procedures.
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
JADE is proposed as a two-layer evaluation framework for agentic AI on open-ended professional tasks to improve stability and align with expert rubrics.
Calculating Mutual Information between a Reward Maximizer and its Environment
A theoretical study quantifying the mutual information between an optimal policy and its environment in controlled Markov processes.
Codex Micro
A discussion regarding Codex Micro, likely focused on its technical specifications or hardware capabilities.
Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM
Thinking Machines has released a massive 975B parameter open-weights LLM, pushing the boundaries of available open-weights models.
Stripe, Advent offer to buy PayPal for more than $53B
Stripe and Advent are reportedly making a bid to acquire PayPal for over $53 billion.
Inkling: Our Open-Weights Model
Introduction of Inkling, a new open-weights model for developers.