AI/ML arXiv cs.AI

PLATO: Pointer Learner for Agent and Task Openness

Presents PLATO, a pointer-network-based actor and GNN critic designed to handle open agent and task spaces in multi-agent reinforcement learning.

AI/ML arXiv cs.AI

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering

Introduces Matryoshka Agent, a hierarchical framework that decomposes complex machine learning engineering tasks into an orchestrator and sub-agents.

AI/ML arXiv cs.AI

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

Identifies failure modes in RL alignment for small language models and proposes a robust system involving adapter reinitialization and precision adjustments.

AI/ML arXiv cs.AI

ScalableRAG: High-Quality RAG at Zero Ingestion Cost

Introduces ScalableRAG, a method for high-quality retrieval-augmented generation that requires zero ingestion cost by using on-the-fly aggregative reasoning.

AI/ML arXiv cs.AI

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Presents DMAPO, a data-centric approach to preference optimization that uses multi-evaluator agreement to select a small, high-confidence dataset.

AI/ML arXiv cs.AI

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

Introduces RRS-10K, a new benchmark for vision-language models (VLMs) specifically focused on rare remote sensing images, especially military-related scenes.

AI/ML arXiv cs.AI

Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings

Presents Aletheia, an offline-first clinical decision support system for low-resource healthcare settings, based on a fine-tuned Qwen2.5-3B-Instruct model using QLoRA.

AI/ML arXiv cs.AI

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

Introduces AdaKP, an online selector that adaptively chooses knowledge-point subsets during reinforcement learning to improve reasoning in LLMs for competition-level mathematics.

AI/ML arXiv cs.AI

MusiChat: Vibe Composing for Music Creation

MusiChat is a conversational music creation system that allows iterative refinement of compositions through a hierarchical controllable music generation framework.

AI/ML arXiv cs.AI

Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation

Analyzes Semantic IDs (SIDs) in generative recommendation and proposes Item-Supported Decoding (ISD) to improve recommendation accuracy without retraining.

AI/ML arXiv cs.AI

Localized Anomaly Detection via Differentiable D-vine Copulas

Proposes a novel estimation framework for D-vine copulas using gradient-based MLE and beam search for more effective and interpretable localized anomaly detection.

AI/ML arXiv cs.AI

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Explores the evidential basis of claims made by multimodal LLMs regarding chart visualizations and suggests the need for systems that distinguish evidence-based claims from interpretations.

AI/ML arXiv cs.AI

SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence

Introduces SAFAARI, a multi-agent framework for NL-to-SQL tasks that improves schema linking in enterprise data environments, along with the SEAL evaluation metric.

AI/ML arXiv cs.AI

CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification

Presents CogEEGAgent, an autonomous agent for EEG analysis grounded in MNE-Python, combining LLM intent interpretation with deterministic scientific validation.

AI/ML arXiv cs.AI

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

Discusses the psychological impacts of conversational AI and proposes design directions to reduce harm and promote user well-being.

Homelab/Self-Hosting Hacker News

More Tailscale tricks for your jailbroken Kindle

A discussion on advanced Tailscale configuration tricks for users with jailbroken Amazon Kindle devices.

Software Engineering Hacker News

Cracking Windows Open: Porting RADV to Win32

An effort to port the RADV (Radeon Vulkan Driver) to Win32, aimed at opening up Windows graphics capabilities.

Software Engineering Hacker News

User Interfaces of the Demo Scene

An exploration of the unique and highly optimized user interfaces used within the demo scene community.

Software Engineering Hacker News

Log is non-monotonic in PHP and Lua

A technical observation regarding the non-monotonic behavior of the log function in PHP and Lua.

AI/ML arXiv cs.AI

On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?

A study evaluating LLMs as alternatives to traditional corpora for specialized terminology translation between English and French.