All Articles
17740 articles total
300k-Year-Old Cave Site Explored in Northern Israel
Exploration of a 300,000-year-old cave site in Northern Israel.
Linux for the Sega MegaDrive
Porting Linux to run on the Sega MegaDrive hardware.
How to corrupt an SQLite database file
A technical exploration of how to intentionally corrupt an SQLite database file.
Alan Kay on the meaning of "object-oriented programming" (2003)
A 2003 talk by Alan Kay discussing the conceptual meaning and origins of object-oriented programming.
ICE Tracks Down Woman to Force Her to Delete Instagram Post
ICE uses tracking to force a woman to delete a post on Instagram.
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Introduces COMPASS, a unified multimodal framework for better composition-intent control in image generation and perception.
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
Presents BV-Blend, a framework to stabilize advantage estimation in critic-free reinforcement learning for LLM alignment.
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance
Analyzes AI governance and adoption using evolutionary game theory and Lean 4 for machine-checking proofs.
TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction
Introduces TrajRS, a framework using randomized smoothing to provide certified robustness for pedestrian trajectory prediction.
ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models
Proposes ComMem, a dual-memory system (fast visual cache and slow textual prototypes) for test-time adaptation of VLMs.
One million passports leaked online
A massive data leak has exposed one million passports online, raising significant privacy and security concerns.
Walter S. Arnold–Sculptor/Stone Carver
A feature on Walter S. Arnold, a sculptor and stone carver.
The AI jobs debate just got messier
A new report suggests that high-intensity AI adoption is actually increasing headcount, including entry-level roles, contradicting the narrative that AI eliminates junior jobs.
Recursive Self-Evolving Agents via Held-Out Selection
Researchers introduce RSEA, a recursive self-evolving agent that improves its own performance across diverse benchmarks by evolving natural-language artifacts using a strict held-out selection gate.
Data and Evaluation Closed-Loop for Model Capability Enhancement
This paper proposes a 'capability slice' framework to create a closed-loop system that maps LLM evaluation failures directly to targeted data interventions during pre-training.
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
GPTNT is a new benchmark for real-time multimodal agent collaboration based on the game 'Keep Talking and Nobody Explodes', highlighting current LLM failures in asynchronous communication.
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
IMCBench is introduced as a multi-turn, image-grounded medical conversation benchmark to evaluate the safety and accuracy of multimodal LLMs in clinical settings.
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories
DynaSteer is a dynamic representation editing framework that steers LLM reasoning trajectories toward truth by monitoring entropy and projecting purified truth vectors.
Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas
VirtueMap provides a framework for profiling the ethical dispositions of LLMs using Aristotelian virtue ethics and a set of non-lethal ethical dilemmas.
An AI agent for treatment reasoning over a biomedical tool universe
ATHENA-R1 is an AI agent trained via reinforcement learning over 212 biomedical tools to perform complex treatment reasoning for FDA-approved drugs.