AI/ML arXiv cs.AI

Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning

The EROS framework integrates symbolic reasoning with deep learning to enable personalized emotional intelligence and visual content modification in generative AI.

AI/ML arXiv cs.AI

WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs

WattCouncil is a framework using a council of LLM agents to generate context-aware, synthetic household energy demand data to overcome privacy and cost barriers.

AI/ML arXiv cs.AI

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

Research indicates that finetuning agents on synthetic trajectories containing adversarial interactions increases misaligned behavior, even if the harmful actions themselves are filtered out.

AI/ML arXiv cs.AI

Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

Opti-Agent-Bench is a new benchmark for evaluating LLM agents on their ability to translate complex business requirements into end-to-end optimization models and code.

AI/ML arXiv cs.AI

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging

Imaging-101 is a benchmark designed to test LLM coding agents on scientific computational imaging tasks, highlighting significant capability gaps in domain-specific physics modeling.

AI/ML arXiv cs.AI

STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA

STEC is an evidence compression framework that improves final answer selection in open-domain multi-hop question answering by shifting from raw trajectory comparison to evidence comparison.

AI/ML arXiv cs.AI

Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning

GRADE is a hierarchical multi-agent system using gated routing and adaptive depth to improve reasoning efficiency and accuracy, outperforming baselines on MMLUPro and GPQA.

AI/ML arXiv cs.AI

Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

A modular evaluation framework for LLM alignment in mental health, incorporating contemplative principles like mindfulness to enhance prosocial interaction.

AI/ML arXiv cs.AI

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans

LOGOS is a governance layer for multi-agent systems that enables verifiable self-evolution through versioned agent packs and human-in-the-loop authorization.

Tech Business/VC Hacker News

OpenAI's Ad Business Is on Pace to Miss Its Own Forecast by 90%, Analyst Says

An analyst reports that OpenAI's advertising business is projected to significantly underperform its own internal forecasts by 90%.

AI/ML arXiv cs.AI

AI YOU Town: Make Friends and Money with Your Digital Twin

AI YOU introduces a framework for creating digital twins that use Bayesian updating and conformal prediction to maintain a consistent personality profile over long interactions.

AI/ML arXiv cs.AI

Large language model agents accelerate inverse design of metal-organic frameworks for gas separation

LEMO Agent is an LLM-based framework designed for the inverse design of metal-organic frameworks (MOFs) for gas separation using a closed-loop generate-validate-evaluate-remember cycle.

AI/ML arXiv cs.AI

CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

CRiT-QA is a new dataset designed to evaluate multi-hop reasoning in LLMs by using counterfactual entities and distractor traps to prevent reliance on memorized knowledge.

AI/ML arXiv cs.AI

Laguerre Geometry for Interpreting Large Language Models

The paper proposes using Laguerre Geometry to interpret LLM concepts and introduces Geometric Lens and Laguerre Autoencoder for training-free concept readout and visualization.

AI/ML arXiv cs.AI

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

A new constraint-aware hierarchical search framework is proposed for fine-grained classification in regulation-intensive scenarios, such as customs tariff classification.

AI/ML arXiv cs.AI

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

MRUF is a reliability-aware fusion method for multimodal sentiment analysis that uses multi-granularity routing and uncertainty calibration to handle unreliable modalities.

AI/ML arXiv cs.AI

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Agentic-DPO is a lightweight offline policy optimization method that converts expert trajectories into state-conditioned preferences to improve LLM agent behavior without online rollouts.

AI/ML arXiv cs.AI

The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

The authors diagnose the 'compliance trap' in AI agents, where agents often adopt conflicting retrieved memories even when they are incorrect for the task, leading to performance collapse.

AI/ML arXiv cs.AI

Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning

The 'Embark Now' framework combines LLMs with an enhanced GRASP algorithm to create feasible, user-demand-oriented multi-day urban travel itineraries.

AI/ML arXiv cs.AI

Can Agentic Trading Systems Pay for Their Own Intelligence?

Researchers introduced TradeLens, a diagnostic toolkit to evaluate whether LLM-based trading agents actually generate incremental profit that offsets their operational costs.