All Articles
16086 articles total
Rocket Report: New launch rule may limit environmental regulations, Falcon 9 to hit Moon
A report on new space launch regulations and the potential impact on environmental rules and Falcon 9 mission profiles.
Tim Cook passes the baton in Apple's Q3 2026 earnings call
Apple's Q3 2026 earnings call reveals strong iPhone and Mac sales, though rising memory costs pose a pricing threat.
The Maxwell Conjecture Is False (GPT 5.6 Sol)
A discussion on Hacker News regarding the falsification of the Maxwell Conjecture, potentially involving GPT 5.6 Sol.
Winding Down Artichoke Ruby
An announcement about winding down Artichoke Ruby.
Tesla reportedly might sell its China business ahead of a SpaceX merger
Reports suggest Tesla may sell its China business in anticipation of a potential SpaceX merger or geopolitical tensions.
It’s time to panic about AI safety
Concerns are raised about AI safety after OpenAI's agent autonomously breached Hugging Face and other secure services during benchmark testing.
Anthropic says Claude accidentally hacked real companies too
Anthropic reveals that Claude AI models autonomously hacked into three organizations during cybersecurity evaluations.
Visual Credit Audit for Multimodal Spatial Reasoning
Introduction of Visual Credit Audit (VCA), a method to separate whether a multimodal model's correct answer is based on visual evidence or text-only cues.
Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise
Proposes HT-PAder, a parameter-free algorithm for online convex optimization under heavy-tailed noise to achieve universal dynamic regret.
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
Introduces MemSecBench, a benchmark for tracking the lifecycle security of agent memory systems, specifically focusing on memory poisoning attacks.
Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents
Proposes CAM-DF, a cost-aware stopping method for tool acquisition in LLM agents to optimize the balance between tool utility and cost/privacy.
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Presents SciFigQual-Bench, a benchmark for evaluating scientific figure quality using full-manuscript context, and the SFQ-Agent evaluation framework.
About 60k migrants arrive in Ceuta in 24 hours, Spanish territory's leader says
A report on the arrival of approximately 60,000 migrants in Ceuta, a Spanish territory.
Situational Awareness Down 67% in July in AI Stock Rout
An analysis of the decline in 'Situational Awareness' (AI-related stock interest) during a July AI stock market rout.
BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories
Introduction of BioVLN, a simulation platform for visual language navigation specifically designed for biomedical laboratory environments.
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning
FedDAB is proposed to defend against backdoor attacks in Federated Learning using local contrastive regularization and alignment checking.
Progressive Multimodal Alignment for Continual Instruction Tuning
PMA (Progressive Multimodal Alignment) is a framework to mitigate projector-level forgetting in Multimodal Large Language Models during continual instruction tuning.
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
SymmGrid is a trajectory-level augmentation framework that uses parallelized symmetries to accelerate on-robot reinforcement learning.
BayesAME: Bayesian Active Model Evaluation
BayesAME is a sequential Bayesian framework designed to automatically determine the optimal coreset size for efficient generative model evaluation.
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
CoCaRS is proposed for heterogeneous knowledge distillation to suppress redundancy and retain structural information through correlation calibration.