All Articles
17730 articles total
The operating cost starts after the demo
A Hacker News discussion focusing on the hidden operational costs that emerge after the initial successful demonstration of a technology.
Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters
Research indicating that the effectiveness of Chain-of-Thought prompting in LLMs depends on the actual reasoning content and validation steps rather than mere verbosity.
Relevance Is Not Permission: Warranted Attention for Value Contributions
Introduction of 'Warrant', a path-localized interface that improves model prediction by ensuring attention relevance is translated into actual evidence.
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars
FacePlex is a unified streaming framework for joint real-time speech and facial motion generation for conversational avatars.
MirrorCode: AI can rebuild entire programs from behavior alone
MirrorCode introduces a long-horizon benchmark where AI agents must reimplement entire software projects based solely on observed behavior.
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
Dynamo is a training-free framework that enables Vision-Language Models to evolve reusable reasoning skills and executable visual tools without weight updates.
From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
Research on spiking agents showing that 'self-caused credit' allows for the development of a durable behavioral self and prevents forgetting during task learning.
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
AIDA is a domain adaptation framework for visual RL that uses 'adaptive imagination' to augment scarce target data for better sim-to-real transfer.
The Many-Body Problem of the Data Centre
A philosophical exploration of the data center as the 'body' of AI and its relationship with human desire and capital.
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
EvalSafetyGap provides a conceptual framework and audit to address failures in LLM evaluation and AI safety measurement.
Crypto exchange OKX wants AI agents to hire and pay each other
Crypto exchange OKX is creating a marketplace for AI agents that integrates payments, identity, and reputation to enable agents to hire and pay one another.
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
Researchers propose a Faithful Warm-Start (FWS) strategy to improve the visual grounding and stability of Vision-Language Models during reinforcement learning.
AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills
AlgoSkill treats algorithm design as a sequential decision-making process using a library of typed skills and Monte Carlo Tree Search for verification-guided refinement.
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
The ACPO framework introduces a decentralized decomposition of the joint policy gradient to improve cooperative tasks in Multi-Agent Reinforcement Learning.
SAT-RTS: A systematic framework for tactical knowledge extraction and visualization-based analysis in real-time strategy games
SAT-RTS is a framework for extracting and visualizing tactical knowledge from real-time strategy games using a cluster-centric BK-tree algorithm.
Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts
HRL-IM/CBS combines influence map hashing and cluster-based scripts within a hierarchical RL framework to improve StarCraft micromanagement.
Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model
A study compares temporal feature extractors in EEG foundation models, finding that pretrained time-series models like MOMENT can be effective frozen extractors.
Propagation of~Interval Belief Structures and~Imprecise Copulas for~Neural Network Verification
This research presents a sound framework for the quantitative verification of neural networks using interval belief structures and imprecise copulas to handle uncertainty.
Structural Certification for Reliable Physical Design with Language Models
The PHACT framework ensures reliable physical design by separating the LLM's proposal stage from a deterministic certification engine.
Open Problems in Constitutional Preference Reconstruction
Researchers identify key problems in Constitutional Preference Reconstruction and propose ICAI+ to improve the agreement between constitution-executor systems.