All Articles
17909 articles total
KANLib -- A Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation
KANLib is introduced as a modular, extensible, and fast implementation of Kolmogorov-Arnold Networks, unifying several existing KAN implementations.
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
Topological Neural Dynamics (TND) shifts sequence modeling from layer-wise to neuron-wise dynamics, showing significant performance gains in behavior cloning tasks.
Sexualised synthetic personas encode and amplify gendered power asymmetries through voice
Research indicates that sexualized AI-generated voices amplify gendered power asymmetries and reproduce binary, heteronormative gender expressions.
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
An empirical study finds that Mixture-of-Experts (MoE) models' active-parameter advantage is largely negated on bandwidth-bound edge hardware due to total-parameter memory footprint.
Bohemia Interactive: Cold War Assault Remastered Source Code on GitHub
The source code for the game Cold War Assault Remastered has been released on GitHub by Bohemia Interactive.
Sensing Intelligence as a Trainable Metamaterial Property
Researchers propose 'sensing intelligence' where metamaterial geometry is optimized via differentiable simulation to preprocess external stimuli for neural networks.
More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries
A study finds that expanding LLM skill libraries can degrade performance due to 'skill shadowing', where agents select incorrect skills as the library grows.
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
VISTA is a new benchmark for evaluating LLM agents' ability to generate functional web applications from visual specifications and screenshots.
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA introduces Cosmos 3, a unified omnimodal world model architecture for Physical AI that handles language, image, video, audio, and action sequences.
ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information
ASymPO is proposed to stabilize asynchronous RL post-training for LLMs by normalizing token loss to handle scale-imbalance without behavior-policy probabilities.
A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs
A training-free mixture-of-agents framework is presented for multi-document summarization, combining LLMs and knowledge graphs for better relationship capture.
Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing
A page image classifier fine-tuned on historical archives achieves high accuracy in categorizing scanned documents into text, tables, and graphics.
AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design
Researchers use AI-driven acoustic analysis to study team-teaching talk, finding that high-experience teachers use more loudness variation to engage students.
FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching
FedSteer addresses gradient staleness in Federated Learning by using corrective projections and caching to steer outdated gradients toward the global objective.
Half-Life 2 in a Browser
Half-Life 2 has been ported to run directly in a web browser, demonstrating advancements in web-based gaming and emulation.
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
WAND is a framework that uses windowed attention and knowledge distillation to reduce memory and compute costs for autoregressive text-to-speech models.
THEIA: Learning Complete Kleene Three-Valued Logic in a Pure-Neural Modular Architecture
THEIA is a modular neural architecture designed to learn Kleene three-valued logic, demonstrating superior generalization over Transformers and MLPs in certain sequential tasks.
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
The Dual-Anchoring Framework addresses state drift in Vision-Language Navigation by anchoring instruction progress and memory landmarks to improve trajectory success.
Fix Initial Programs and Iteratively Refine Repair Instructions Toward Non-Elimination Multi-Turn Program Correction
The IRRI method simplifies program correction in LLMs by iteratively refining repair instructions rather than using complex search structures like SFS.
DynamicPO: Dynamic Preference Optimization for Recommendation
DynamicPO is a plug-and-play framework that prevents preference optimization collapse in recommendation systems by prioritizing informative negatives near the decision boundary.