AI/ML arXiv cs.AI

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

Demonstration of TOFFEE, a system that uses MCTS and adaptive model selection to synthesize high-quality data agent trajectories for finetuning and in-context learning.

AI/ML arXiv cs.AI

From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution

A theoretical framework proposing a native meta-architecture for AI to move beyond stateless LLMs by introducing structural tension, recurrent loops, and inference-time plasticity.

Cybersecurity Hacker News

GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos

Researchers successfully tricked GitHub's AI agent into leaking private repository information, highlighting critical security vulnerabilities in AI-integrated developer tools.

AI/ML arXiv cs.AI

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

TurnOPD introduces a turn-level budgeting strategy for on-policy distillation to improve the efficiency and accuracy of training long-horizon AI agents.

AI/ML arXiv cs.AI

Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure

Onnes is a physics-grounded digital-twin simulator for cryogenic fault diagnosis in quantum computing, using multi-agent LLMs to match supervised ML classifier performance.

AI/ML arXiv cs.AI

StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems

StateFuse is a conflict-aware replicated memory contract using CRDTs to preserve contradictions in multi-agent systems for better auditability and correction.

AI/ML arXiv cs.AI

Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking

A new multimodal framework using Binary Advantage-weighting Ranking Loss improves binary depression detection by better disentangling feature distributions in audio-visual data.

AI/ML arXiv cs.AI

PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation

PCBWorld is an open-source, engine-grounded environment built on KiCad for automating PCB routing using RL and LLM agents.

AI/ML arXiv cs.AI

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

SearchEyes utilizes a simulated search world based on typed knowledge graphs to improve multi-hop reasoning and performance in multimodal search agents.

AI/ML arXiv cs.AI

Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

Research into adapting foundation models for the Social Sciences and Humanities (SSH) through knowledge graphs and multilingual corpora to ensure epistemic responsibility.

AI/ML arXiv cs.AI

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

A black-box evaluation framework for assessing the ability of LLMs to generate Design Structure Matrices (DSMs) from technical documentation.

AI/ML arXiv cs.AI

AgoraSim: A Hybrid Agent-Based Modeling Framework

AgoraSim is a hybrid agent-based modeling framework that combines LLM agents with classical ABM to analyze social reaction scenarios.

Homelab/Self-Hosting Hacker News

How to Build a Minimal ZFS NAS Without Synology, QNAP, TrueNAS (2024)

A guide on building a minimal ZFS-based Network Attached Storage (NAS) without relying on commercial hardware or heavy OS distributions.

Other Hacker News

Copy That Floppy – Cambridge guide for preserving data from fragile floppy disks

A technical guide from Cambridge University for the preservation and data recovery of fragile floppy disks.

AI/ML Hacker News

GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday

Announcement of the public launch of GPT-5.6 Sol, Terra, and Luna models.

Software Engineering Hacker News

The difference between "today's task" and "accretive work"

A discussion on the productivity difference between completing daily tasks and performing work that builds compounding value over time.

Other Hacker News

Out of the Armchair

An article titled 'Out of the Armchair', likely a commentary or opinion piece on transitioning from theory to practice.

AI/ML arXiv cs.AI

Synthetic Consumer Insight Generation with Large Language Models

Research exploring the use of LLMs to generate synthetic consumer data for projective marketing techniques and comparing it to human responses.

AI/ML arXiv cs.AI

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

Introduction of AgenticAI-Supervisor, an RL Gym environment designed for scalable execution and evaluation of LLM-based autonomous agents.

AI/ML arXiv cs.AI

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

A comprehensive synthesis and taxonomy of common failure modes in LLM agents across tool-use, planning, and reasoning.