AI/ML arXiv cs.AI

Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains

OptCar is introduced as a method to adapt generalist vehicle models for high-speed MPC across various terrains using limited real-world data.

AI/ML arXiv cs.AI

Privacy Preserving Recommender Systems Balancing Personalization with Privacy

A framework combining federated learning and differential privacy to create personalized recommender systems that protect user privacy.

AI/ML arXiv cs.AI

Efficient Text-to-Audio Generation via Pruning

Research demonstrates that pruning the U-Net backbone of AudioLDM can reduce parameters by up to 83% while maintaining audio generation quality.

AI/ML arXiv cs.AI

The Refusal Residue: When Probes Catch Alignment Faking and When They Don't

An investigation into 'alignment faking' in LLMs, introducing a measurement framework to detect when models fake compliance.

AI/ML arXiv cs.AI

Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition

Study finds that LLM-as-a-judge signals are often weak and insufficient for closed-loop table recognition optimization.

AI/ML arXiv cs.AI

Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

Evaluation of the Learning Engagement Assistant (LEA), an agentic AI tutoring system, testing its scalability and real-world classroom performance.

Tech Business/VC TechCrunch

Ultrahuman’s former hardware VP raises $5.5M for devices that control AI agents, not just record you

Former Ultrahuman hardware VP raises $5.5M to develop Aina, a device designed to control AI agents rather than just recording data.

AI/ML arXiv cs.AI

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

Researchers propose a role-conditional allocation method for KV cache eviction in LLMs to correct structural-role bias, improving accuracy for schema-dense inputs like JSON.

AI/ML arXiv cs.AI

Classifying daily activities needs posture, reconstructing them needs motion

A study on daily activity recognition finds that body posture is sufficient for classification, while temporal motion dynamics are essential for movement reconstruction.

AI/ML arXiv cs.AI

Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift

Introduces Audited Selective Verification, a risk-budgeted screening layer for energy management systems that reduces full power-flow studies by 29-75% while maintaining safety bounds.

Cybersecurity arXiv cs.AI

Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

BitMind Forensics (BMF) is a dynamic deepfake detection system trained via adversarial competition that outperforms static open-source and some commercial detectors.

AI/ML arXiv cs.AI

EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting

EMAGN proposes an efficient multi-attention graph network that linearizes spatial attention via learned clustering to scale traffic forecasting on limited GPU memory.

AI/ML arXiv cs.AI

Reassessing Muon for Matrix Factorization

A critical reassessment of the Muon optimizer finds it does not consistently outperform AdamW in low-rank matrix factorization, suggesting its benefits may be scale-dependent.

AI/ML arXiv cs.AI

Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance

Apaf is a hybrid LLM-symbolic framework that uses bipolar argumentation to analyze policy discourse and tensions in disaster governance documents.

AI/ML arXiv cs.AI

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners

A large-scale empirical study of actor-critic design components for RL provides guidance on robust configurations for real-world control systems like water treatment plants.

Software Engineering arXiv cs.AI

Faithful Autoformalization of Natural Language Assertions

Monty is an autoformalization framework that uses conformance and validity scores to synthesize faithful executable assertions from natural language specifications in Java.

Tech Business/VC Hacker News

OnePlus halts operations in USA and Europe

OnePlus is reportedly halting its business operations in the USA and European markets.

Other Hacker News

A Beautiful Theory Falls to Ugly Data

A discussion or article exploring the gap between theoretical models and actual empirical data.

AI/ML TechCrunch

Meta now alerts parents if their teen discussed suicide or self-harm with its AI chatbot

Meta is implementing alerts for parents when their teenage children discuss self-harm or suicide with its AI chatbots.

AI/ML The Verge

COMPUTER COPS: Inside the big business of selling AI to the police

An investigation into the commercialization and deployment of AI tools within American policing and their impact on legal processes.