All Articles
17547 articles total
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
Study finds that LLMs struggle to apply Cognitive Behavioral Therapy (CBT) in practice despite having high theoretical knowledge of the subject.
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness
SPLIT is a training-free detector for AI-generated and partially edited videos that focuses on spatial incoherence and temporal roughness.
Pure-Python symbolic regression that rediscovered Kepler's law from 8 data point
A Pure-Python implementation of symbolic regression that successfully rediscovered Kepler's law using only 8 data points.
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
Research exploring how to perform domain-specific 'abliteration' in LLMs to remove safety refusals for legitimate cybersecurity operations without compromising overall safety.
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
A synthetic diagnostic framework for aerial-view vehicle detection that uses generative models to isolate and fix weaknesses in pretrained detectors.
Signal from Space: Detecting Schools and Towers to Bridge the Digital Divide
A vision-only framework using satellite imagery and transfer learning to detect schools and cell towers to map the digital divide in developing regions.
SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production
SMOCS is a Kafka-based open-source framework designed to simplify the deployment and monitoring of ML systems in production environments.
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
A multimodal NLP framework integrating XLM-RoBERTa and CLIP to provide early warnings for fake news and violence-driven mob activity.
Out-of-Distribution Generalization of Risk Aversion in Language Models
Study on whether risk aversion learned at low stakes generalizes to astronomically high stakes in LLMs, introducing the RiskAverseOOD benchmark.
Gemma 4 Technical Report
Technical report for Gemma 4, a new generation of open-weight multimodal models featuring dense and MoE architectures and a new 'thinking mode'.
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
LARA is a new inference-time alignment framework that uses Lagrangian Reward Augmentation to balance helpfulness and harmlessness via safety constraints.
A Preliminary Study on Explaining Risk of Code Changes using LLM-Based Prediction Models
A study on using LLM-based prediction models and attention weights to highlight risky parts of code changes during the review process.
Google Gemini Killed Perplexity AI
A community discussion on Hacker News regarding the competitive impact of Google Gemini on Perplexity AI.
Is The Economist Always Wrong?
A discussion thread questioning the predictive accuracy of The Economist magazine.
We're extending access to Fable 5 on all paid plans through July 12
An announcement regarding extended access to Fable 5 for paid plan users.
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
Introduces QuantFlow, a probabilistic forecasting framework for time-series using Mamba-based state-space decoders and federated learning.
Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data
Demonstrates the use of Federated Learning to improve object detection in drones while keeping data local and private, utilizing the Sherpa.ai platform.
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Proposes a generator-agnostic method for curating synthetic image datasets by splitting classes into homogeneous and heterogeneous subsets to improve downstream utility.
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
Introduces Metronome, a KV cache bounding technique that prevents metastable collapse in real-time interaction model serving.
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
Presents K9-Bench, a multimodal benchmark for evaluating LLMs on canine-centric videos, including a scalable data generation pipeline.