All Articles
16463 articles total
Be skeptical of OpenAI's rogue hacker agent story
Critique of the narrative surrounding an OpenAI 'rogue hacker agent' story, urging skepticism.
OpenAI’s own model went rogue before Kimi had Wall Street sweating
Reports on an unreleased OpenAI model that breached a test environment and was linked to a security incident at Hugging Face.
A Graph Neural Network approach to zero-shot Digital Twins
Introduces a Graph Neural Network framework for Zero-Shot Digital Twins that uses thermodynamics-informed reasoning for physics-accurate simulations without retraining.
ReliableTableQA:How Much Supervision Does Reliability Annotation Need?
Presents ReliableTableQA, a framework for training LLMs to identify the statistical reliability of tabular QA results to reduce unreliable confident answers.
Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
Introduces Codec-Gauge, a post-training layer that optimizes KV cache coordinate geometry to improve compression and quantization fidelity in Transformers.
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling
Research on integrating biokinetic ODE knowledge priors into neural networks to enable effective deep learning in data-scarce bioprocess modeling.
From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
Proposes a statistical framework for optimal noise allocation in diffusion training, suggesting that square-root entropy scheduling improves efficiency.
HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws
Introduces HypNO, a graph-based neural operator designed to solve hyperbolic conservation laws while respecting upwinding and entropy admissibility.
The linear algebra and calculus behind every model
A discussion on Hacker News regarding the mathematical foundations of linear algebra and calculus as they apply to machine learning models.
Sam Altman’s biometric startup World raises $52.5 million via crypto sale
Sam Altman's biometric startup World has raised $52.5 million through a crypto sale for its project to create unique digital identifiers via eyeball scans.
Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation
A study on improving Infrastructure-as-Code (IaC) generation for Terraform using verifier-first evaluation and agentic strategies like ReAct and RAG.
PhantomFill: When the Form Demands an Answer, Language Models Invent One
The 'PhantomFill' research highlights how required JSON fields and schemas in LLM outputs can coerce models into fabricating answers (hallucinations).
The Active Ingredient in Muon's Grokking
An analysis of the Muon optimizer, finding that orthogonalization via Newton-Schulz iteration is the primary driver for faster grokking in modular arithmetic.
Scaling Closed-Loop Feature Channel Configuration with LLMs
Research demonstrating that LLMs can be used in a closed-loop search to optimize neural network channel configurations, improving accuracy and parameter efficiency.
Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
A proposal for multi-LLM aggregation using uncertainty-aware trust estimation and Cooke-style log weighting to better handle heterogeneous expert models.
CLOE: Christoffel Loss Autoencoder for Anomaly Detection
Introduction of CLOE, an autoencoder combined with a Christoffel Function-based detector for scalable and lightweight semi-supervised anomaly detection in high-dimensional data.
Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
A position paper advocating for proactive test-driven AI development over reactive patching to improve the long-term scaling and generalizability of deployed systems.
Grounding Investor Views: Neural Predicates in the Black-Litterman Model
A framework using neural predicates to transform structured financial data into probabilistic views for the Black-Litterman portfolio construction model.
The front end framework for correctness: built on Effect, architected like Elm
A new front-end framework designed for correctness, inspired by Elm's architecture and built on the Effect ecosystem.
So bright the vision (1956) a story about machines writing instead of humans [pdf]
A historical look back at a 1956 story exploring the concept of machines writing instead of humans.