Filtered out
Friday · May 29, 2026
94Filtered stories
94Raw rejects
94Unique rejects
source cap reached: 25below stricter editorial threshold: 22A library for rendering SVG from Markdown.: 1AI research on low-data learning, relevant for ML engineers.: 1AI research on VLM spatial planning capabilities; relevant for future AI applications.: 1Applies multimodal foundation models and RL for zero-shot control; relevant for AI developers.: 1Benchmark for LLM agents in strategic games.: 1Benchmark for LLM financial verification, relevant to applied research.: 1Benchmark for LLM peer review behavior, relevant to AI research.: 1Benchmark for LLM social reasoning in multi-agent settings.: 1Benchmark for traffic forecasting, relevant to ML engineers.: 1Edge-deployable LLM for trajectory prediction; niche for robotics devs.: 1Education survey, not directly relevant to developers.: 1Evaluates LLM reasoning, relevant for AI-assisted development.: 1Explains dense retrieval for AI search systems.: 1Health text generation with LLMs, relevant for applied AI devs.: 1Improves automated survey generation for developers using AI.: 1Improves lightweight GUI agents for mobile automation.: 1Improves LLM agent efficiency, reducing latency/cost.: 1Improves RAG for document QA, relevant to AI devs: 1Leverages LLM agents for complex scientific reasoning, demonstrating new application areas.: 1LLM multi-agent framework for storytelling, tangential to dev work.: 1LLM optimization for devs, but niche research.: 1LLM robustness research relevant to prompt engineering.: 1LLM safety robustness via optimizer choice, relevant for AI developers.: 1Materials science benchmark, not directly for developers.: 1Medical AI paper, not relevant to software development.: 1Niche forestry robotics, not relevant to software dev.: 1Niche RL research; limited direct dev impact.: 1Optimizes LLM quantization for generative tasks, improving deployment efficiency.: 1PEFT for diffusion LLMs; niche but relevant to AI developers.: 1Regulatory compliance QA benchmark for LLMs, tangential to dev.: 1Reinforcement learning paper, tangential to most developers.: 1Relevant for AI developers working on LLM reasoning and decoding.: 1Reproducibility metadata format for ML evaluations.: 1Robotic world model benchmark, tangential to dev work.: 1Specialized BCI research, not general dev tooling.: 1Specialized EEG research, not general dev work.: 1Specialized energy forecasting, not general dev work.: 1Specialized molecular dynamics, not general dev.: 1Tangential educational AI research, not directly applicable.: 1Tangential to applied ML, not for general devs.: 1Tangential to developers; focuses on clinical trial trends.: 1Tangential to developers; symbolic planning research.: 1Theoretical bandit paper, niche for ML researchers.: 1Theoretical game theory, not directly applicable to daily dev work.: 1Theoretical MARL paper, not directly applicable to daily dev work.: 1Theoretical ML paper, not directly applicable to daily dev work.: 1Theoretical RL paper, not directly applicable to daily dev work.: 1
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systemsFHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Series Inpainting and ForecastingOn the Geometry of Games and their SolversPractitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey EvidenceDifferentiable Belief-based Opponent ShapingTrends in AI and Human-AI Interaction in Clinical Trials -- A Hybrid Human-AI ExplorationSurfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teachers and Students Using AIEvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular DynamicsQuantifying and Optimizing Simplicity via Polynomial RepresentationsOmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science SubfieldsMind-Omni: A Unified Multi-Task Framework for Brain-Vision-Language Modeling via Discrete DiffusionBehavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference PredictionOrthogonal Concept Erasure for Diffusion ModelsBridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution SemanticsMiraBench: Evaluating Action-Conditioned Reliability in Robotic World ModelsImproving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language ModelsLLM-Evolved Domain-Independent Heuristics for Symbolic AI PlanningBitTP: The Lightweight Trajectory Prediction Model with BitLLM for Edge-DevicesUncertainty-Aware Transfer Learning for Cross-Building Energy Forecasting: Toward Robust and Scalable District-Level Energy ManagementCertified Policy Optimisation for Nested Causal Bandits via PAC-Bayes RiskOptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based DistillationBehavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy PredictionBenchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Modelsmarkdown-svg-rendererCitation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question AnsweringFrom XXLTraffic to EvoXXLTraffic: Scaling Traffic Forecasting to Sensor-Evolving NetworksFinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement VerificationHarmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic SchedulingWhen and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming LoopConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE CompressionPassNet: Scaling Large Language Models for Graph Compiler Pass GenerationBattery-Sim-Agent: Leveraging LLM-Agent for Inverse Battery Parameter EstimationThink Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured Health Text GenerationPTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game?NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMsMEMENTO: Leveraging Web as a Learning Signal for Low-Data DomainsPRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted ReviewingHarnessing non-adversarial robustness in large language modelsDiagnosing Harmful Continuation in Answer-Correct Long-CoT Training TracesPaper Agents, Paper Gains: An Empirical Analysis of DeFi Investment AgentsTailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility🔬ESM: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHubAligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order OptimizationThe Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language ModelingEntropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language ModelsDenseSteer: Steering Small Language Models towards Dense Math ReasoningFrontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypesThe Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion ModelsXetrieval: Mechanistically Explaining Dense RetrievalUI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI AgentsPlanning with the Views via Scene Self-ExplorationCroissant Tasks: A Metadata Format for Reproducible Machine Learning EvaluationsRobust and Efficient Guardrails with Latent ReasoningCoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool RetrievalThe Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial PressureBEAMS: Benchmarking and Evaluating AI for Modeling and SimulationBetter Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction CorrectionReasonOps: Operator Segmentation for LLM Reasoning TracesRethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground TruthVFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element AnalysisReview Arcade: On the Human Alignment and Gameability of LLM ReviewsRubric-Guided Process Reward for Stepwise Model RoutingReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal ControlMINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMsDeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey GenerationHiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringTRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT EvaluationLFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMsSAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic SearchIndexing the Unreadable: LLM-Native Recursive Construction and Search of Service TaxonomiesArchitecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR BenchmarkPRO-CUA: Process-Reward Optimization for Computer Use AgentsReliable Reasoning with Large Language Models via Preference-Based Maximum SatisfiabilityBeyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph ModelingAdopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the WildBeyond Attack Success Rate: Temporal Logit Observability for LLM Safety FailuresNotation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI SystemsThe Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data PlaneBeyond Consensus: Trace-Level Synthesis in Mixture of AgentsNICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMsGRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM AgentsThe Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIFHallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic CachingMoment-KV: Momentum-Based Decode-Time KV Cache Compression for Long GenerationRedundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent TrajectoriesGTA: Generating Long-Horizon Tasks for Web Agents at ScaleAgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and SecurityBenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM AgentsSkillsInjector: Dynamic Skill Context Construction for LLM AgentsOpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution TrajectoriesDeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement LearningWhen Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMsGoverning Technical Debt in Agentic AI SystemsGPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activity Chain Generation