AI/ML arXiv cs.AI

EviDAG: Auditable Causal DAG Authoring with Biomedical Literature

EviDAG is a browser-based system that helps authors create auditable causal directed acyclic graphs (DAGs) using evidence from biomedical literature via LLMs.

AI/ML arXiv cs.AI

MemTX: Transactional Belief Commit for Stateful Agent Memory

MemTX introduces a transactional belief-commit protocol for stateful agent memory to prevent irreversible actions based on polluted or stale updates.

AI/ML arXiv cs.AI

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

EviBack is a reinforcement learning approach for search agents that uses evidence-constrained teacher backoff to improve RAG systems.

AI/ML arXiv cs.AI

Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

A new study uses LLM-based simulation with 'Perceived Transaction Cost' personas to better model human responses to energy policy interventions.

AI/ML arXiv cs.AI

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

SciConsolidate is a method for converting verified runtime experience in scientific computing into transferable procedural knowledge to improve LLM capabilities.

AI/ML arXiv cs.AI

"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research

An interview study explores the collaborative practices and sociotechnical challenges of interdisciplinary AI and healthcare research.

AI/ML arXiv cs.AI

FFNet: MetaMixer-based Efficient Convolutional Mixer Design

FFNet proposes an efficient convolutional mixer design (MetaMixer) that replaces self-attention with large kernel convolutions while retaining the query-key-value framework.

AI/ML arXiv cs.AI

Real-time Spatial Retrieval Augmented Generation for Urban Environments

A proposed real-time spatial RAG architecture for urban environments implemented via the FIWARE ecosystem to improve smart city digital twins.

AI/ML arXiv cs.AI

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

Researchers introduce a conceptual framework and synthetic dataset created in NVIDIA Omniverse to train VLMs for visual perspective taking in robots.

AI/ML arXiv cs.AI

On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy

A systematic review and taxonomy of human-centered explainable AI (XAI) systems, focusing on evaluation metrics for both AI novices and data experts.

AI/ML arXiv cs.AI

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering

Introduces SAE-Steering, a method using Sparse Autoencoders to control reasoning strategies in Large Reasoning Models (LRMs).

AI/ML arXiv cs.AI

JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI

JobMatchAI is a production-ready job matching platform using knowledge graphs, semantic search, and the JobSearch-XS benchmark.

AI/ML arXiv cs.AI

DSevolve: Enabling Real-Time Adaptive Scheduling on Dynamic Flexible Job Shop with LLM-Evolved Heuristic Portfolios

DSevolve is a framework for real-time adaptive scheduling in job shops using LLM-evolved heuristic portfolios and a neural selector.

AI/ML arXiv cs.AI

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem

A theoretical exploration of AI alignment, arguing for an autonomy-supporting 'parenting' approach as AGI attains subject status.

AI/ML arXiv cs.AI

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

A study on Medical VQA revealing that grounding-by-cropping often degrades performance and proposing a hybrid fine-tuning approach to restore localization.

AI/ML arXiv cs.AI

The Scaling Properties of Implicit Deductive Reasoning in Transformers

Research investigating the scaling properties of implicit deductive reasoning in Transformers using Horn clauses.

AI/ML arXiv cs.AI

AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading

AlphaCrafter is a multi-agent framework for quantitative trading that replaces loose natural-language workflows with structured policy specifications.

AI/ML Hacker News

A pharmacy chain in Vermont implemented AI for efficiency

A pharmacy chain in Vermont is utilizing AI to improve operational efficiency.

Other Hacker News

Angels in Coptic Magic I: Introduction

An introductory piece on 'Angels in Coptic Magic', discussing historical or mystical texts.

AI/ML arXiv cs.AI

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

MemLens is a value-aware memory management system for LLM-based agents that uses an interactive dashboard for analytics and memory lifecycle visualization.