All Articles
17650 articles total
Is the iPhone birth control? Causal evidence from AT&T's 2007-2011 monopoly [pdf]
A research paper examining the causal relationship between the introduction of the iPhone and birth rates during AT&T's monopoly period.
EFF letter to FTC on X consent order [pdf]
An EFF letter to the FTC regarding a consent order involving X (formerly Twitter).
Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
Introduction of Procedural Memory Distillation (PMD), a method to help LLMs retain procedural knowledge from training trajectories into their weights.
The Agentic Garden of Forking Paths
A study on how AI agents can replicate human ideological gaps in data analysis and the introduction of the 'm-value' and 'Agentic Bootstrap' to detect selective reporting.
Janus: a Playground for User-Involved Agentic Permission Management
Introduction of Janus, a playground system for evaluating and implementing user-involved permission management for autonomous AI agents.
Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
Proposes Semi-CoT, a framework for semi-supervised chain-of-thought learning that uses unlabeled questions to create pseudo-reasoning supervision.
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration
Presents OPINE-World, an LLM agent that learns programmatic world models online from interaction using ontology error for exploration.
Scaling Trends for Lie Detector Oversight in Preference Learning
Explores the scaling trends of Lie Detector Oversight (SOLiD) in preference learning to monitor and prevent deceptive behavior in LLMs.
EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation
A three-agent LLM pipeline designed for generating scientific hypotheses for Earth observation grounded in NASA's knowledge graph.
Perform DFU Restores on Apple Silicon Macs with Macvdmtool (2021)
A discussion on using Macvdmtool to perform DFU restores on Apple Silicon Macs.
PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations
PACE is a neuro-symbolic framework that combines neural predictive models with symbolic reasoning to generate plausible and actionable counterfactual explanations in AI.
Auto-FL-Research: Agentic Search for Federated Learning Algorithms
Auto-FL-Research (AFR) introduces a constrained coding-agent workflow to automate the search for optimal federated learning algorithms.
The Wiola Architecture for Efficient Small Language Models
The Wiola architecture is a novel Small Language Model (SLM) featuring Spiral Rotary Positional Encoding and Gated Cross-Layer Attention, designed for efficiency.
Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases
Agent4cs is a multi-agent framework designed to summarize large, hierarchical codebases using specialized agents for summarization, keyword extraction, and quality assurance.
When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations
Proposes a difficulty-routed service-control architecture for customer-service agents to escalate complex requests while maintaining efficiency for routine tasks.
CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
CreativityNeuro uses contrastive weight steering to improve divergent thinking and reduce mode collapse in LLMs without requiring fine-tuning.
Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Adapts a mixture-of-experts diffusion language model (DiffusionGemma-26B) for radiology report drafting, offering faster decoding and superior infill capabilities over autoregressive models.
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
Explores using Reinforcement Learning with Verifiable Rewards (RLVR) to improve tool-use agents performing enterprise SaaS workflows in Atlassian environments.
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Introduces MedAgentBench-v3 to diagnose and improve RL in clinical environments, identifying capability and format-knowledge barriers in clinical agents.
Ask HN: Looking for work, donations, friendship, or advisors
A community thread on Hacker News where individuals seek employment, mentorship, or social connections.