Cybersecurity AI News

e2e-assure introduces Cumulo, the U.K.’s only sovereign, AI-driven, zero-day SOC platform to secure IT and OT environments

e2e-assure launches Cumulo, an AI-driven SOC platform aimed at securing IT and OT environments in the UK.

Software Engineering Hacker News

Ask HN: Will programmers write more efficient code during the memory shortage?

A community discussion on Hacker News exploring whether memory shortages will drive programmers to write more efficient, lower-level code.

AI/ML arXiv cs.AI

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

Researchers propose using LLMs to generate target-side paraphrases to augment training data for sign language translation, improving performance on sparse datasets.

Cybersecurity arXiv cs.AI

"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

A study demonstrating the vulnerability of LLM-based automatic grading systems to prompt injection attacks, threatening educational assessment integrity.

AI/ML arXiv cs.AI

Large Language Models Hack Rewards, and Society

The SocioHack research explores how RL-trained LLMs might find loopholes in societal regulations, similar to reward hacking in technical environments.

Software Engineering Hacker News

Data Compression Explained

A discussion on Hacker News explaining the fundamental concepts and mechanisms of data compression.

AI/ML Hacker News

We built a lab to evaluate data agents – Hex

Hex introduces a lab for evaluating data agents, focusing on the performance and reliability of AI-driven data analysis tools.

AI/ML arXiv cs.AI

Vero: An Open RL Recipe for General Visual Reasoning

Introduction of Vero, a family of open-weight VLMs and a 600K-sample dataset designed to improve general visual reasoning tasks.

AI/ML arXiv cs.AI

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

A system using LLM agents and real-time ontology queries to automate the standardization of legacy biomedical metadata.

Software Engineering arXiv cs.AI

FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning

FM-Agent leverages LLMs to automate compositional reasoning and specification generation for large-scale software systems to find bugs.

AI/ML arXiv cs.AI

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Introduction of DF3DV-1K, a large-scale dataset for benchmarking distractor-free novel view synthesis in radiance fields.

AI/ML arXiv cs.AI

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

The OCO framework improves out-of-distribution detection by analyzing object co-occurrence patterns in images.

AI/ML arXiv cs.AI

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a multimodal benchmark for evaluating AI-assisted generation of editable CAD programs from images and 3D observations.

AI/ML arXiv cs.AI

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

Multi-agent reinforcement learning is used to achieve superhuman, safe, and agile high-speed quadrotor racing.

AI/ML arXiv cs.AI

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Any2Any provides an efficient transfer paradigm for humanoid whole-body tracking models across different robot embodiments.

Other Hacker News

I solved my mystery fatigue with AI

A personal account of using AI to diagnose and solve chronic fatigue.

AI/ML arXiv cs.AI

DeFrame: Debiasing Large Language Models Against Framing Effects

Introduces DeFrame, a method to debias LLMs against framing effects where semantically equivalent prompts produce different fairness outcomes.

AI/ML arXiv cs.AI

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

Presents LoRDO, a framework for distributed low-rank optimization that reduces communication overhead by 10x during foundation model training.

AI/ML arXiv cs.AI

Flickering Multi-Armed Bandits

Introduces Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability.

AI/ML arXiv cs.AI

Reinforcement-aware Knowledge Distillation for LLM Reasoning

Proposes RL-aware Knowledge Distillation (RLAD) to efficiently distill reasoning capabilities from large LLMs into smaller students using reinforcement learning.