AI/ML arXiv cs.AI

Lifting Embodied World Models for Planning and Control

A method for 'lifting' embodied world models by mapping high-level actions to low-level joint sequences, reducing search complexity for humanoid robot control.

Hardware/Chips arXiv cs.AI

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

Presentation of PostEDA-Bench, a hierarchical benchmark for evaluating LLM agents in the final stages of circuit design, focusing on DRC fixing and PPA convergence.

AI/ML arXiv cs.AI

GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding

Introduction of Group-Query Latent Attention (GQLA), a modification of MLA that enables hardware-adaptive decoding paths for different GPU architectures without retraining.

Other arXiv cs.AI

Global Automation Atlas

A large-scale study using LLMs to map automation exposure across 124 economies, analyzing how task content and country-level conditions affect labor displacement.

AI/ML arXiv cs.AI

Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning

Tunable MAGMAX is proposed as a preference-aware model merging framework to mitigate catastrophic forgetting in continual learning by adjusting task-specific performance.

Other arXiv cs.AI

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

A theoretical exploration of the 'Cognitive Kardashev Scale', quantifying the maximum possible machine cognition based on energy supply and hardware efficiency.

Open Source Hacker News

Codeberg Bans Cryptocurrency Projects

Codeberg, an open-source git hosting service, has implemented a ban on projects related to cryptocurrency.

AI/ML Hacker News

Run large language models at home, BitTorrent‑style

A discussion on running large language models (LLMs) locally using a BitTorrent-style peer-to-peer distribution method.

Software Engineering Hacker News

NixOS is more complicated than you think but OpenCode fixed it

Discussion regarding the complexity of NixOS and how a project called OpenCode aims to simplify its usage.

AI/ML arXiv cs.AI

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines

Research identifying a 'selection bottleneck' in multi-agent LLM pipelines, suggesting that selector quality is more critical than model diversity for output quality.

AI/ML arXiv cs.AI

Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

Introduction of Doctorina MedBench, a dialogue-based evaluation framework for agent-based medical AI using simulated physician-patient interactions.

AI/ML arXiv cs.AI

M-RAG: Semantic Key-Value Indexing for Retrieval-Augmented Generation

Proposed M-RAG, a semantic key-value indexing layer for RAG that separates retrieval keys from generation payloads to optimize token budgets.

AI/ML arXiv cs.AI

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification

FVRuleLearner introduces an Operator Reasoning Tree (Op-Tree) to automate formal verification and improve the generation of SystemVerilog Assertions.

AI/ML arXiv cs.AI

Robust Reasoning Benchmark

The Robust Reasoning Benchmark (RRB) reveals that open-weights LLMs are highly susceptible to textual perturbations and 'Intra-Query Attention Dilution'.

AI/ML arXiv cs.AI

CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations

CPGRec+ is a balance-oriented framework for personalized video game recommendations utilizing LLMs and GNNs to improve accuracy and diversity.

AI/ML arXiv cs.AI

Why Do Vision Language Models Struggle To Recognize Human Emotions?

Study analyzes why Vision Language Models struggle with human emotion recognition, citing head-class bias and inability to handle dense temporal sequences.

AI/ML arXiv cs.AI

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction

A new framework called SKETCH improves long-horizon vessel trajectory prediction by using semantic key-points to maintain global directional consistency.

AI/ML arXiv cs.AI

Toward Learning POMDPs Beyond Full-Rank Actions and State Observability

Researchers propose a method to learn Partially Observable Markov Decision Processes (POMDPs) using spectral approaches and tensor decomposition to reason about systems with hidden states.

AI/ML arXiv cs.AI

Training and Simulation of Quadrupedal Robot in Adaptive Stair Climbing and Descending for Indoor Firefighting: An End-to-End Reinforcement Learning Approach

A two-stage reinforcement learning approach allows quadruped robots (Unitree Go2) to adaptively climb and descend diverse indoor staircases for firefighting missions.

AI/ML arXiv cs.AI

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

LinguistAgent is a reflective multi-model platform designed to automate complex linguistic annotation tasks like metaphor identification.