All Articles
18306 articles total
e2e-assure introduces Cumulo, the U.K.’s only sovereign, AI-driven, zero-day SOC platform to secure IT and OT environments
e2e-assure launches Cumulo, an AI-driven SOC platform aimed at securing IT and OT environments in the UK.
Ask HN: Will programmers write more efficient code during the memory shortage?
A community discussion on Hacker News exploring whether memory shortages will drive programmers to write more efficient, lower-level code.
Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models
Researchers propose using LLMs to generate target-side paraphrases to augment training data for sign language translation, improving performance on sparse datasets.
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems
A study demonstrating the vulnerability of LLM-based automatic grading systems to prompt injection attacks, threatening educational assessment integrity.
Large Language Models Hack Rewards, and Society
The SocioHack research explores how RL-trained LLMs might find loopholes in societal regulations, similar to reward hacking in technical environments.
Data Compression Explained
A discussion on Hacker News explaining the fundamental concepts and mechanisms of data compression.
We built a lab to evaluate data agents – Hex
Hex introduces a lab for evaluating data agents, focusing on the performance and reliability of AI-driven data analysis tools.
Vero: An Open RL Recipe for General Visual Reasoning
Introduction of Vero, a family of open-weight VLMs and a 600K-sample dataset designed to improve general visual reasoning tasks.
Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent
A system using LLM agents and real-time ontology queries to automate the standardization of legacy biomedical metadata.
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
FM-Agent leverages LLMs to automate compositional reasoning and specification generation for large-scale software systems to find bugs.
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Introduction of DF3DV-1K, a large-scale dataset for benchmarking distractor-free novel view synthesis in radiance fields.
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
The OCO framework improves out-of-distribution detection by analyzing object co-occurrence patterns in images.
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
CADBench is a multimodal benchmark for evaluating AI-assisted generation of editable CAD programs from images and 3D observations.
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
Multi-agent reinforcement learning is used to achieve superhuman, safe, and agile high-speed quadrotor racing.
Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
Any2Any provides an efficient transfer paradigm for humanoid whole-body tracking models across different robot embodiments.
I solved my mystery fatigue with AI
A personal account of using AI to diagnose and solve chronic fatigue.
DeFrame: Debiasing Large Language Models Against Framing Effects
Introduces DeFrame, a method to debias LLMs against framing effects where semantically equivalent prompts produce different fairness outcomes.
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Presents LoRDO, a framework for distributed low-rank optimization that reduces communication overhead by 10x during foundation model training.
Flickering Multi-Armed Bandits
Introduces Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability.
Reinforcement-aware Knowledge Distillation for LLM Reasoning
Proposes RL-aware Knowledge Distillation (RLAD) to efficiently distill reasoning capabilities from large LLMs into smaller students using reinforcement learning.