All Articles
17780 articles total
Waymo and Uber quietly part ways in Phoenix
Waymo and Uber have ended their three-year partnership in Phoenix, Arizona.
Chronic Kidney Disease Prognosis Prediction Using Transformer
Researchers introduced ProQ-BERT, a transformer-based framework for predicting Chronic Kidney Disease progression using multi-modal electronic health records.
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
The SMFA framework and S-MLLMUn Bench are proposed to enable selective unlearning of privacy-sensitive information in multimodal LLMs without degrading general capabilities.
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Health-ORSC-Bench is a new large-scale benchmark to measure over-refusal and safe completion quality of LLMs in healthcare contexts.
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
The GAIA system introduces a data flywheel to train intuitive critic models that improve the test-time scaling performance of GUI agents.
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
The JustAsk framework demonstrates that autonomous code agents can be used to systematically recover hidden system prompts from frontier LLMs.
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
CausalFlip is a new benchmark designed to test whether LLMs are performing true causal reasoning or merely relying on semantic matching.
Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning
COffeE-PSRO is introduced as a conservative equilibrium discovery method for offline game-theoretic multiagent reinforcement learning.
SEA-TS: Self-Evolving Agent for Autonomous Code Generation of Time Series Forecasting Algorithms
SEATS is a self-evolving agent framework that autonomously generates and optimizes code for time series forecasting algorithms.
Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report
A new policy-iteration algorithm, iPI, is presented to decide the safety of states in non-deterministic problems with guaranteed polynomial worst-case runtime.
Wallace the 6 inch f/2.8 telescope, building it, and hiking with it
A detailed account of building and using a custom 6-inch f/2.8 telescope, focusing on the process and practical application.
The Radiation Exposure Lie
A discussion questioning the common narratives around radiation exposure.
Ornith-1.0: self-improving open-source models for agentic coding
Introduction of Ornith-1.0, a set of open-source models designed for self-improving agentic coding.
Anthropic and Gov. Newsom forge deal allowing California government to use Claude at half price
Anthropic has reached an agreement with the California government to provide Claude AI services at a discounted rate.
South Korean tech giants commit over $550B to ease ‘ RAMageddon’
South Korean memory chip giants commit $550 billion to expand fabrication capacity to address AI-driven RAM shortages.
At $499, Apple’s M3-powered iPad Air is a good deal
A review and deal highlight for the M3-powered iPad Air, noting its performance gains and value at a discounted price.
"Generate" the Future of Work through AI: Empirical Evidence from Online Labor Markets
An empirical study on how generative AI is displacing labor and inducing skill transitions in online programming-intensive job markets.
Agentic Episodic Control
Presents Agentic Episodic Control (AEC), a new architecture integrating LLMs into reinforcement learning to improve data efficiency and generalization.
Symmetry-Aware Transformer Training for Automated Planning
Introduces a symmetry-aware contrastive learning objective for Transformers to improve their efficiency in automated planning tasks.
PreferThinker: Reasoning-based Personalized Image Preference Assessment
Proposes PreferThinker, a reasoning-based framework for personalized image preference assessment using a predict-then-assess paradigm.