AI/ML arXiv cs.AI

An Ontology for Machine Learning Interatomic Potentials

The MLIPs ontology provides a standardized OWL 2 DL framework to describe machine learning interatomic potentials, improving reproducibility in materials science.

AI/ML arXiv cs.AI

Characterisation of Density-based FM generation methods in the context of Information Fusion

Research on density-based Fuzzy Measure (FM) generation methods to improve information fusion and provide confidence intervals for fusion results.

AI/ML arXiv cs.AI

CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

CAPT is a continuous autoregressive transformer designed for calcium population dynamics that generalizes across different datasets and species.

AI/ML arXiv cs.AI

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

SeekJudge is a practical reward framework using specialized agents to provide accurate and efficient rewards for RL in computer-use agents.

AI/ML arXiv cs.AI

TopoFE: topology-aware LLM-guided Automated Feature Engineering

TopoFE is a topology-aware evolutionary framework that uses LLMs to automate feature engineering for tabular data more efficiently.

Other Hacker News

Does every question mark deserve a Betteridge?

A discussion on the 'Betteridge's law of headlines', suggesting that any question mark in a headline implies a negative answer.

Software Engineering Hacker News

Hooray for the Sockets Interface

A technical appreciation of the Sockets interface and its role in network programming.

Other Hacker News

Beyond Greece and Rome

An exploration of intellectual and cultural history beyond the classical Greek and Roman traditions.

Tech Business/VC Hacker News

Chip stocks slide in US and Asia as AI jitters rattle investors

Chip stocks are declining in the US and Asia due to investor concerns regarding AI growth and sustainability.

AI/ML arXiv cs.AI

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

Introduces ConsistencyGate, a write-time admission control mechanism to prevent hallucinated facts from contaminating the external memory of LLM agents.

AI/ML arXiv cs.AI

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

Presents Reason Popper-ly, a neurosymbolic framework using inductive logic programming to verify and correct in-context reasoning steps in LLMs.

AI/ML arXiv cs.AI

Stress-testing large language model agents in a robotic chemistry laboratory

An evaluation of LLM agents in a physical robotic chemistry lab, revealing significant gaps in long-horizon planning and physical executability.

AI/ML arXiv cs.AI

SymStep: Symbolic Step Verification for Logical Reasoning

Introduces SymStep, a symbolic step verification method that uses a constraint propagator to ensure logical consistency in LLM reasoning chains.

AI/ML arXiv cs.AI

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

Proposes a structured spatio-temporal transformer block that explicitly models interaction types to reduce the need for deep model stacks in multi-entity reasoning.

AI/ML arXiv cs.AI

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

Presents a hierarchical diagnosis framework for optimizing Triton kernels on NPUs by linking runtime symptoms to compiler IR structure.

Software Engineering Hacker News

Multiple Mouse Cursors in Wayland

A discussion regarding the implementation of multiple mouse cursors within the Wayland display server protocol.

AI/ML arXiv cs.AI

Coordinated Networking for On-Device Agent-Augmented Real-Time Communication

Researchers introduce HFS, a framework for on-device agent-augmented real-time communication that optimizes traffic between human video streams and AI agent data flows.

AI/ML arXiv cs.AI

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

A theoretical study on certified runtime safety for AI agents that use tools, focusing on the limits of what can be enforced by runtime guardrails.

AI/ML arXiv cs.AI

Physical AI Governance: From Theory to Practice Across Life Cycle

A comprehensive survey providing a governance framework and a five-stage lifecycle for 'Physical AI' systems (embodied AI) to ensure safety and trustworthiness.

AI/ML arXiv cs.AI

How Well Can AI Generate Backlogs from App Mockups?

An evaluation of how well AI can generate software backlogs from app mockups, finding that hybrid prompting with architectural context improves precision.