AI/ML arXiv cs.AI

SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia

SafeStep is an AI-driven travel assistance system for elderly people with dementia, combining LLMs with behavioral prediction to generate personalized failure scenarios and interventions.

AI/ML arXiv cs.AI

Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications

A study evaluates the effectiveness of transcript-based and tool-based safeguards for speech-to-speech LLM assistants in cars, finding both insufficient due to latency and non-deterministic behavior.

AI/ML arXiv cs.AI

Explaining Weather Bulletins via ILP

This research utilizes Inductive Logic Programming (ILP) via the FastLAS2 framework to generate interpretable hypotheses explaining weather bulletins from raw meteorological data.

AI/ML arXiv cs.AI

Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems

A new method using matrix-based differentiable logic programming is proposed to reduce reasoning shortcuts in neurosymbolic systems, ensuring more robust concept mapping and constraint satisfaction.

Hardware/Chips arXiv cs.AI

Identifying Good Rules for Efficient SAT Encodings of Single-Constant Multiplication Using Machine Learning

A neuro-symbolic framework using Graph Neural Networks is introduced to accelerate SAT encoding for the Single Constant Multiplication problem in hardware design.

Software Engineering arXiv cs.AI

Bound-Founded Semantics for Answer Set Programming with Difference Constraints: Preliminary Report

This report introduces a many-sorted variant of the Bound-founded Logic of Here-and-There (HTb) to provide a unified semantic foundation for Answer Set Programming with linear constraints.

AI/ML arXiv cs.AI

Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents

The paper proposes a cue-anchored working memory model for coding agents to ensure reliability by treating memory as a harness property rather than an agent choice.

AI/ML arXiv cs.AI

Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

This research examines the complex interactions between compression, MoE routing, and quantization in multimodal edge intelligence, arguing they should not be treated as independent optimizations.

AI/ML arXiv cs.AI

GuardianAgentBench: Where Agents Fail and How to Guard Them

The authors introduce GuardianAgentBench, a benchmark to evaluate the safety and reliability of LLM agents across various domains and adversarial attack modes.

AI/ML arXiv cs.AI

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

The paper introduces Workflow-Localized Mechanism Learning (WML) to help agents repair and reuse knowledge for structured procedural skills.

AI/ML arXiv cs.AI

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Naju is a native discrete state-space model designed for long-sequence memory tracking by decoupling retention and writing mechanisms.

AI/ML arXiv cs.AI

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

This study investigates the trustworthiness and stability of LLM-generated summaries by proposing a two-level diagnostic protocol.

AI/ML arXiv cs.AI

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

EmoAgent-R1 uses reinforcement learning and dynamic agent specialization to improve multimodal emotion understanding in MLLMs.

Homelab/Self-Hosting arXiv cs.AI

HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices

HiMe is a privacy-first, locally deployable platform for real-time health monitoring using wearable device data and LLM agents.

AI/ML arXiv cs.AI

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

The researchers present Faster IndexTTS-2, an acceleration method for autoregressive text-to-speech models using NVIDIA TensorRT and TensorRT-LLM.

AI/ML arXiv cs.AI

Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation

This paper analyzes the limitations of current semantic-ID based generative recommendation systems in handling cold-start items.

AI/ML Hacker News

Claude Cookbook

Anthropic provides a 'Cookbook' of examples and guides for developers using Claude.

AI/ML arXiv cs.AI

Code Monitor Red Teaming for Public-Test-Passing Code

Researchers introduce CodeMonitorBench to test if LLM verifiers can identify bugs in code that has already passed public tests.

AI/ML arXiv cs.AI

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

The MisKnow-Agent framework reveals that Deep Research LLM agents are vulnerable to misleading knowledge, often adopting false conclusions in final reports.

AI/ML arXiv cs.AI

Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning

A new source-prior-driven selective adaptation method for diffusion models improves the trade-off between target-specific generation and the retention of general capabilities.