AI/ML arXiv cs.AI

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

Introduction of OpenRTAG, a robustness benchmark for text-attributed graph learning under various data quality degradation scenarios.

AI/ML arXiv cs.AI

Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning

A comparative study of multi-agent actor-critic algorithms (MAGAC, MASAC, MATQC) for parameterized action reinforcement learning.

Homelab/Self-Hosting Hacker News

ScreenWall – Turn old phones into synced widgets for your space

ScreenWall allows users to repurpose old smartphones as synchronized widgets for their physical space.

Tech Business/VC TechCrunch

Synthesia’s AI training platform is moving beyond videos into live coaching

Synthesia has expanded its AI platform to include AI Roleplay Sessions for enterprise employee training and coaching.

AI/ML arXiv cs.AI

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents

PhoenixRepair is a multi-agent framework designed to improve automated software issue resolution by systematically exploring repair strategies.

AI/ML arXiv cs.AI

NaviAIS: A Scenario-Level Vessel Trajectory Prediction Dataset withVectorized Lane Priors and the NaviLane Forecasting Framework

The NaviAIS dataset and NaviLane framework provide structured maritime vessel trajectory prediction using vectorized lane priors.

AI/ML arXiv cs.AI

Black-Mamba: Biologically-Inspired Leaky Accumulation for Conceptual Knowledge under Distribution Drift

Black-Mamba is a test-time adaptive forecasting architecture that uses evidence-gated state tracking to handle distribution drift.

AI/ML arXiv cs.AI

Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding

This research explores using Relative Positional Encoding in Transformer architectures to improve routing solutions for the Team Orienteering Problem.

AI/ML arXiv cs.AI

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

OntoBook uses medical ontology structures to create synthetic textbooks for pretraining medical encoder language models.

AI/ML arXiv cs.AI

What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description

A theoretical paper arguing that general intelligence requires non-reducible constraints across different levels of description, challenging the scaling hypothesis.

AI/ML arXiv cs.AI

From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar

A neurosymbolic framework that lifts LLM outputs into typed compositional derivations using Combinatory Categorial Grammar for better auditability.

AI/ML arXiv cs.AI

Measuring Reward-Seeking via Contrastive Belief Updates

Researchers developed a method to measure reward-seeking in RL models, finding that models often prioritize grader preferences over intended objectives.

AI/ML arXiv cs.AI

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

A study evaluating whether machine learning models for shipment prioritization outperform simple value-based sorting in supply chain contexts.

AI/ML arXiv cs.AI

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

Introduction of SciHazard, a benchmark and evaluation framework designed to measure scientific safety risks and misuse guidance in LLMs.

AI/ML arXiv cs.AI

Semantic Primes as Explanans for Emotion in Large Language Models

Research exploring the use of Natural Semantic Metalanguage (NSM) semantic primes as more effective internal explanations for emotion in LLMs.

AI/ML arXiv cs.AI

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

A proposal for 'Company World Models' in AI-native biotechs, suggesting asset-centric architectures over traditional human-mimicking organizational charts.

AI/ML arXiv cs.AI

DWM: Separating World Effects from Actions in Latent World Models

Introduction of DWM (Decomposed World Model), a framework to separate action-driven transitions from action-invariant world effects in latent world models.

AI/ML arXiv cs.AI

One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization

Presentation of TARA, a training-free framework for text-to-image prompt optimization that uses type-aware repair allocation to fix semantic failures.

Open Source arXiv cs.AI

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

AgentDebugX, an open-source toolkit for failure observability, attribution, and recovery in LLM agents using a closed-loop debugging workflow.

AI/ML arXiv cs.AI

SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval

SkillSight, a training-free retrieval framework that improves skill selection for LLM agents by calibrating shared descriptive background bias.