All Articles
16443 articles total
Waymo reportedly mulling a breakup with Uber
Reports that Waymo is considering ending its contractual partnership with Uber.
Operational Identity: A Finite Audit of Declared and Implemented Rules of Sameness
A formal study on 'Operational Identity', analyzing the divergence between declared and implemented rules of sameness in record systems.
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning
Presentation of GPE, a benchmark and framework to evaluate LLM robustness against generative engine optimization (GEO) poisoning in fact verification.
Self-Supervised Bio-Inspired Robotic Trajectory Planning with Obstacle Avoidance
Research on a bio-inspired self-supervised learning framework for robotic trajectory planning with obstacle avoidance.
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Introduction of IssueTrojanBench, which reveals significant security vulnerabilities in AI coding agents like Cursor and Claude Code when facing malicious issue requests.
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
An audit of diversity metrics in LLM ensembles, finding that many metrics track capability rather than actual diversity.
Emergent Compositional Skills in Mixture-of-Experts VLAs
Research on Mixture-of-Experts (MoE) VLAs that emergently learn compositional robot policies without pre-specified task decomposition.
HARP: The Human--AI Research Platform
Introduction of HARP, a platform designed for systematic research into Human-AI Interaction by providing controlled mock scenarios with live agents.
Frontier Financial Judgement: Can agents tell what might move a stock?
Introduction of the Frontier Financial Judgement benchmark to evaluate AI agents' ability to replicate expert human financial analysis and stock movement predictions.
Scaling Interpretable Transformers with Parity Bottleneck Layers
The ParityTransformer introduces Deep Parity Bottlenecks to create language models that are interpretable by design rather than relying on post-hoc sparse autoencoders.
SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
SalesLoop is a reinforcement learning framework for CRM lead ranking that uses performance-aware rewards and Discriminative GRPO to align model predictions with real-world outcomes.
Adaptive Multi-Horizon Reinforcement Learning
A new multi-horizon reinforcement learning approach that adaptively selects temporal horizons to improve decision-making in changing environments without manual discount-factor tuning.
From Agent Failures to Text Policies: What Works and What Breaks
Analysis of TextGrad for agent optimization, finding a gap between the ability to follow textual policies and the ability to generate them from experience.
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
Introduction of Spatially Grounded Concept Bottleneck Models (SG-CBM) to improve the interpretability and trustworthiness of AI-assisted breast ultrasound diagnosis.
DS@GT ARC at ImageCLEFmed GANs 2026: Geometric Filtering for Privacy-Preserving CT Slice Generation
A privacy-preserving framework for synthetic lung CT slice generation using Optimal Transport Conditional Flow Matching and geometric filtering to reduce patient re-identification.
A Framework for Reputation Aware Uninorm-driven Consensus Algorithms for Blockchain Networks
A proposed reputation-aware consensus algorithm for blockchain networks using intuitionistic fuzzy sets and uninorm aggregation to improve fairness and inclusivity.
U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation
U-CFR is an inference-time framework for interactive image segmentation that uses uncertainty-guided pseudo-clicks to autonomously refine segmentation masks.
Transition-Related Potentials as Markers of Narrative Comprehension in Continuous EEG
Research demonstrating that narrative context in films leaves measurable signatures in continuous EEG responses, which can be detected using deep neural networks.
Marimo Now Runs in PyCharm
Marimo, a reactive notebook for Python, now has integration and support for running within the PyCharm IDE.
Volkswagen engineers charged with insider trading tied to Rivian joint venture
Volkswagen engineers have been charged with insider trading involving a joint venture with Rivian.