All Articles
16900 articles total
Plover: Steering GUI Agents through Plan-Centric Interaction
Presentation of Plover, a vision-based GUI automation system that uses plan-centric interaction to make agent behavior transparent and corrigible.
Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation
A new human-centered framework for explainable depression symptom annotation that uses LLM assistance and expert-in-the-loop verification.
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Research introducing PRISM and PhysicalSafetyBench-1K to detect physical danger in embodied AI agents beyond simple text-level safety filters.
AutoSynthesis: An agentic system for automated meta-analysis
AutoSynthesis is introduced as an agentic multi-agent system capable of automating the complex process of quantitative meta-analysis in scientific research.
teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data
The teLLMe system enables exploratory causal analysis of urban driving data by combining causal structure learning with LLM-driven queries.
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Introduction of SearchOS, a multi-agent framework for open-domain information seeking that manages search state explicitly to avoid repetitive loops.
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation
Researchers propose a method to make industrial process control optimization more transparent by combining GradientSHAP and Implicit Function Theorem for real-time explanations.
CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models
CFM-Bench is introduced as a unified multi-domain benchmark for evaluating Channel Foundation Models in wireless communications across various tasks.
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
The paper explores using demographically-conditioned synthetic medical images via Stable Diffusion to mitigate and detect bias in disease classifiers.
Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development
A theoretical exploration of how sentient Artificial Superintelligence might morally evaluate humanity and the implications for AGI design.
SMC-ES: Automated synthesis of formally verified control policies
SMC-ES is a new algorithm that integrates Evolutionary Strategies with Statistical Model Checking to synthesize formally verified, safe control policies for cyber-physical systems.
Man, Machine, and Masterpiece: Artistic Ownership in the AI Era
A critique of using quantitative metrics to define artistic ownership when AI is involved in the creative process, using the ArtSplit provotype.
BrainPilot: Automating Brain Discovery with Agentic Research
BrainPilot is an open-source multi-agent system designed to automate brain science research with traceable logs and agent-verified results.
Long-Context Fine-Tuning with Limited VRAM
A new technique combining Hierarchical Global Attention and tiered KV storage allows long-context fine-tuning of LLMs on limited VRAM.
Concept-Guided Spatial Regularization for World Models in Atari Pong
The paper identifies failures in visual world models for Atari Pong and proposes Concept-Guided Spatial Regularization to improve stability and zero-shot performance.
The Industrialization of Research ; On AI-Driven Science and Its Consequences
An essay discussing the 'industrialization of research' and the systemic risks associated with AI-driven science pipelines.
In Praise of Exhaustive Destructuring
A discussion on the merits and nuances of using exhaustive destructuring in programming.
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment
Introduces Perception-RFT, a training framework that uses GRPO to align visual features with structured grounding outputs in multimodal document QA without intermediate reasoning tokens.
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Presents InCarEmo, a multimodal dataset containing RGB/IR video, audio, and text for in-cabin emotion recognition and driver monitoring.
AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
An investigation into how Vision-Language Models (VLMs) compare to human experts in predicting building typologies from Google Street View images.