All Articles
17720 articles total
This could be our best look yet at Samsung’s new wide foldable
Leaked case designs suggest a new wide-style foldable design for the upcoming Samsung Galaxy Z Fold 8.
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Clarus is introduced as a collaboration infrastructure designed to coordinate autonomous research agents for web-scale scientific collaboration.
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
The paper proposes Inoculation Adapters (IA), using LoRAs to suppress undesired traits in AI models more effectively than inoculation prompting.
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
EMPATH is a new multilingual auditor-judge benchmark for evaluating the safety of emotional-support chatbots.
PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning
PromptGNN-sim is a bi-directional fusion framework that integrates Graph Neural Networks (GNNs) and LLMs for improved text-attributed graph learning.
Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering
Rhetor is a multi-agent system that automates live product demonstrations by analyzing a web app and its source code to produce synchronized narration and Q&A.
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
ManimAgent is a self-evolving multimodal agent that uses an episodic memory bank to improve its ability to generate mathematical animations using the Manim library.
BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery
BayesEvolve introduces a belief-guided discovery framework that uses explicit uncertainty-aware belief states to improve autonomous scientific discovery.
Sequential Fairness Auditing with Limited Output Access
A new sequential generalized likelihood-ratio framework is proposed for auditing AI fairness under limited model output access.
Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data
Researchers provide a provable statement that well-calibrated LLMs can serve as low-cost statistical estimators for human-response data in social and behavioral sciences.
The operating cost starts after the demo
A Hacker News discussion focusing on the hidden operational costs that emerge after the initial successful demonstration of a technology.
Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters
Research indicating that the effectiveness of Chain-of-Thought prompting in LLMs depends on the actual reasoning content and validation steps rather than mere verbosity.
Relevance Is Not Permission: Warranted Attention for Value Contributions
Introduction of 'Warrant', a path-localized interface that improves model prediction by ensuring attention relevance is translated into actual evidence.
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars
FacePlex is a unified streaming framework for joint real-time speech and facial motion generation for conversational avatars.
MirrorCode: AI can rebuild entire programs from behavior alone
MirrorCode introduces a long-horizon benchmark where AI agents must reimplement entire software projects based solely on observed behavior.
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
Dynamo is a training-free framework that enables Vision-Language Models to evolve reusable reasoning skills and executable visual tools without weight updates.
From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
Research on spiking agents showing that 'self-caused credit' allows for the development of a durable behavioral self and prevents forgetting during task learning.
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
AIDA is a domain adaptation framework for visual RL that uses 'adaptive imagination' to augment scarce target data for better sim-to-real transfer.
The Many-Body Problem of the Data Centre
A philosophical exploration of the data center as the 'body' of AI and its relationship with human desire and capital.
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
EvalSafetyGap provides a conceptual framework and audit to address failures in LLM evaluation and AI safety measurement.