All Articles
17919 articles total
Grok Build 0.1: Intelligence, Performance and Price Analysis
An analysis of the intelligence, performance, and pricing of Grok Build 0.1.
Bob Iger’s Disney wanted Apple, Twitter, and 007
Former Disney CEO Bob Iger reveals failed attempts to acquire Apple, Twitter, and the James Bond franchise.
Assessing Distribution Shift in Human Activity Recognition for Domain Generalization
A study on distribution shift in Human Activity Recognition (HAR) introducing a new open-source benchmark platform and datasets.
Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching
Introduction of Bidirectional Conditional Flow Matching (Bi-CFM) to solve inverse problems in chaotic systems, showing significant speedups and accuracy gains.
Difference-Making without Making a Difference
A philosophical critique of seven definitions of actual causation, arguing that distinctions between them are meaningless.
Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment
An evaluation of multi-turn LLM dialogues for assessing Non-Functional Requirements (NFRs) in HIPAA compliance, finding low accuracy against expert ground truth.
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
Analysis of automated grading strategies for agentic data analysis systems, introducing a three-layer human-AI grading cascade.
Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models
Introduction of the Match Task to Objective (MTO) framework to optimize fine-tuning and prompt-tuning for encoder-decoder language models.
World Models in Pieces: Structural Certification for General Agents
Proposed structural certification framework to localize reliable transitions for long-horizon planning in general agents' world models.
Vector Graphics in Lil
A discussion about implementing vector graphics in the Lil programming language.
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Themis is an explainable AI framework for Reinforcement Learning with Human Feedback (RLHF), supporting over 200 environments to improve transparency and alignment.
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
SAFARI introduces a tool-augmented diagnostic loop and short-term memory to identify faults in long-horizon agentic trajectories beyond LLM context limits.
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
CineCap is a framework for cinematographic video captioning using structured reasoning and spatio-temporal anchors, outperforming existing baselines.
LaGO: Latent Action Guidance for Online Reinforcement Learning
LaGO uses pretrained LLMs as latent action priors to guide online policy optimization in reinforcement learning, improving success rates in control benchmarks.
Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation
A new branch-and-bound algorithm for constructing cost-optimal decision diagrams for stochastic Boolean function evaluation, proven to be #P-hard.
Decentralised AI Training and Inference with BlockTrain
BlockTrain is a decentralized training protocol that partitions models into independently trainable blocks to enable AI training without centralized GPU clusters.
Scaling Laws for Task-Specific LLM Distillation
The paper derives empirical scaling laws for task-specific LLM distillation and introduces the FinHeadlineMix dataset for domain-specific compression.
Can Scale Save Us From Plasticity Loss in Large Language Models?
Research indicates that large Transformer models eventually lose plasticity—the ability to learn new information—regardless of model size or training setting.
BluTrain: A C++/CUDA Framework for AI Systems
BluTrain is a lightweight C++/CUDA framework for AI training that outperforms PyTorch in throughput and memory efficiency for GPT-2 baselines.
GitHub Is Becoming a Giant AI Code Dump
A community discussion on Hacker News regarding the trend of GitHub becoming a repository for low-quality AI-generated code.