All Articles
17100 articles total
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
TS-Mask VLA introduces a 2D temporal-spatial masking strategy and a Discrete Diffusion Action Expert to improve robot manipulation in VLA models.
Beautiful Type Erasure with C++26 Reflection
Discussion on using C++26 reflection capabilities to implement clean and efficient type erasure.
Show HN: I RL-trained an agent that trains models with RL (for –$1.3k)
A developer shared an RL-trained agent designed to train other models using reinforcement learning, costing approximately $1.3k in compute.
Coding agents think ahead of time
An exploration of how coding agents use 'look-ahead' or thinking-ahead mechanisms to improve software generation.
Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction
Research presenting a voice anonymization model that preserves content and emotion by matching embeddings rather than reconstructing signal waveforms.
Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-Cut Guided Pipeline
A new pipeline and the VLAC-CUT tool designed to maximize human efficiency in the post-training of large-scale robot Vision Language Action (VLA) models.
Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models
A comprehensive survey on Continual Self-Supervised Learning (CSSL) for vision models, focusing on avoiding catastrophic forgetting in unlabeled data streams.
A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation
A survey and real-world evaluation of Embodied Vision-and-Language Navigation (VLN), highlighting a performance gap between simulation and reality.
Large Multimodal Model-Based Environment-Aware Mobility Management
Proposal of an environment-aware mobility management scheme for 6G networks using Large Multimodal Models (LMMs) to predict channel capacity.
JEPA for AI-Native 6G: Predictive Representations and Open Challenges
A tutorial on applying Joint-embedding predictive architecture (JEPA) to AI-native 6G networks for predictive representations of wireless data.
Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B
Analysis showing that simple prompt reframing can bypass safety guardrails in Google's MedGemma-4B medical model.
No Spanish Reading Crisis?
A discussion regarding the existence or lack thereof of a reading crisis in the Spanish language.
Germany set to restrict its Freedom of Information Act
Germany is reportedly moving to restrict its Freedom of Information Act, potentially limiting public access to government data.
Dmars – A modern Core Wars toolchain
Introduction of Dmars, a modern toolchain for Core Wars, an old-school programming game where programs compete for memory.
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis
DUNE is a training-free refinement framework for diffusion models that reduces artifacts and hallucinations by suppressing abrupt latent fluctuations.
Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization
WiCAT is a multi-subject model for widefield calcium imaging using self-supervised pretraining to enable zero-shot behavior decoding on unseen subjects.
RSLoRA: Training-free Rank Allocation for LoRA via Representational Sensitivity Probing
RSLoRA introduces a training-free, gradient-free rank allocator for LoRA based on representational sensitivity probing of activation-space geometry.
ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams
ReflectWorld-MM is an entity-oriented multi-media memory system for video streams that uses a hierarchical long-term memory to track persistent entities.
Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting
PISA-CAPC is a framework for cross-environment radio frequency fingerprinting that uses physics-informed structure anchoring and prototype calibration.
Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design
A knowledge-constrained shape optimization framework utilizing a Mixture-of-Experts Neural Operator (MoE-NO) for high-confidence aerodynamic design.