AI 主题导航
按研究方向浏览论文解读,快速进入相关主题与延伸阅读。
AI Agent 127 篇
- Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- PPTArena: A Benchmark for Agentic PowerPoint Editing
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Memory in the Age of AI Agents
- MemEvolve: Meta-Evolution of Agent Memory Systems
- Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
- DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- A Multi-Agent Framework for Stateful Inference-Time Search
- Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
- What's the next frontier for Data-centric AI? Data Savvy Agents
查看其余 115 篇
- What Limits Agentic Systems Efficiency?
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Virtual Agent Economies
- VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Tree Search for LLM Agent Reinforcement Learning
- Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning
- Towards a Science of Scaling Agent Systems
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- The Rise and Potential of Large Language Model Based Agents: A Survey
- The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
- The FM Agent
- The Era of Agentic Organization: Learning to Organize with Language Models
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
- Staircase Streaming for Low-Latency Multi-Agent Inference
- SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
- Scaling Environments for LLM Agents in the Era of Learning from Interaction: A Survey
- ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
- Retrieval Augmented Generation (RAG) for Fintech: Agentic Design and Evaluation
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
- Reinforcement Learning for Machine Learning Engineering Agents
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Re4: Scientific Computing Agent with Rewriting, Resolution, Review and Revision
- QAgent: A modular Search Agent with Interactive Query Understanding
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
- Online Process Reward Leanring for Agentic Reinforcement Learning
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
- MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
- Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
- Mathematical Framing for Different Agent Strategies
- MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning
- MAPEX: A Multi-Agent Pipeline for Keyphrase Extraction
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- LLM imesMapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
- LLM-in-Sandbox Elicits General Agentic Intelligence
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- Kimi K2.5: Visual Agentic Intelligence
- Internalizing World Models via Self-Play Finetuning for Agentic RL
- Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond
- Inefficiencies of Meta Agents for Agent Design
- In-Context Distillation with Self-Consistency Cascades: A Simple, Training-Free Way to Reduce LLM Agent Costs
- How Far Are We from Genuinely Useful Deep Research Agents?
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems
- General Agentic Memory Via Deep Research
- GEM: A Gym for Agentic LLMs
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
- EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
- EvoClaw: Evaluating AI Agents on Continuous Software Evolution
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Effective context engineering for AI agents
- Dynamic Speculative Agent Planning
- Dynamic Affective Memory Management for Personalized LLM Agents
- DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- Beyond Pipelines: A Survey of the Paradigm Shift toward Model-Native Agentic AI
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Are Agents Just Automata? On the Formal Equivalence Between Agentic AI and the Chomsky Hierarchy
- Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
- An Information Theoretic Perspective on Agentic System Design
- Alita-G: Self-Evolving Generative Agent for Agent Generation
- AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
- AI Agent Systems: Architectures, Applications, and Evaluation
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
- Agentic Software Engineering: Foundational Pillars and a Research Roadmap
- Agentic Meta-Orchestrator for Multi-task Copilots
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
- Agentic Inequality
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Adaptation of Agentic AI
- A Survey on Large Language Model based Autonomous Agents
- A Survey on Agentic Multimodal Large Language Models
- A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
- A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
RAG与知识系统 61 篇
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Memory in the Age of AI Agents
- MemEvolve: Meta-Evolution of Agent Memory Systems
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- Tackling the Inherent Difficulty of Noise Filtering in RAG
- Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
- SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
- RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers
查看其余 49 篇
- RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
- Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
- Retrieval Augmented Generation (RAG) for Fintech: Agentic Design and Evaluation
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
- Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
- On the Theoretical Limitations of Embedding-Based Retrieval
- MoM: Mixtures of Scenario-Aware Document Memories for Retrieval-Augmented Generation Systems
- MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
- MeSH: Memory-as-State-Highways for Recursive Transformers
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
- MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
- LLM-guided Hierarchical Retrieval
- LLM-empowered knowledge graph construction: A survey
- Less LLM, More Documents: Searching for Improved RAG
- Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- ImageBind: One Embedding Space To Bind Them All
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
- HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- General Agentic Memory Via Deep Research
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- Fast-weight Product Key Memory
- EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Dynamic Affective Memory Management for Personalized LLM Agents
- DoPE: Denoising Rotary Position Embedding
- Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval
- DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
- CRoPE: Efficient Parametrization of Rotary Positional Embedding
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- Continual Learning via Sparse Memory Finetuning
- CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
- CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
- Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
- CAMformer: Associative Memory is All You Need
- BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval
- AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
推理与强化学习 120 篇
- ΔL Normalization: Rethink Loss Aggregation in RLVR
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
- What is the objective of reasoning with reinforcement learning?
- Wait, Wait, Wait... Why Do Reasoning Models Loop?
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
查看其余 108 篇
- Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- TreeWriter: AI-Assisted Hierarchical Planning and Writing for Long-Form Documents
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- Tree Search for LLM Agent Reinforcement Learning
- Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning
- The Universal Landscape of Human Reasoning
- The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
- The Path Not Taken: RLVR Provably Learns Off the Principals
- The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
- The Illusion of Insight in Reasoning Models
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Statistical Reinforcement Learning in the Real World: A Survey of Challenges and Future Directions
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- Scaling Latent Reasoning via Looped Language Models
- Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
- ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
- Reinforcement learning
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMs
- Reinforcement Learning for Machine Learning Engineering Agents
- Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
- Reflexion: Language Agents with Verbal Reinforcement Learning
- QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
- Prompt Repetition Improves Non-Reasoning LLMs
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
- Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
- Outcome-based Exploration for LLM Reasoning
- Online Process Reward Leanring for Agentic Reinforcement Learning
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- On GRPO Collapse in Search-R1: The Lazy Likelihood-Displacement Death Spiral
- Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning
- LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
- LIMO: Less is More for Reasoning
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- FlowRL: Matching Reward Distributions for LLM Reasoning
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- Exploration v.s. Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Evaluating Parameter Efficient Methods for RLVR
- Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
- Dynamic Speculative Agent Planning
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- DocReward: A Document Reward Model for Structuring and Stylizing
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- Deep Self-Evolving Reasoning
- Data-Efficient RLVR via Off-Policy Influence Guidance
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
- CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
- CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
- BroRL: Scaling Reinforcement Learning via Broadened Exploration
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
- Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- A Survey on Parallel Reasoning
- A Survey of Reinforcement Learning for Large Reasoning Models
- A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms
- A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- A Survey of Inductive Reasoning for Large Language Models
- A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- A Multiobjective Reinforcement Learning Framework for Microgrid Energy Management
模型训练与优化 107 篇
- Weight-sparse transformers have interpretable circuits
- Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
- Supervised learning pays attention
- Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
- Kimi Linear: An Expressive, Efficient Attention Architecture
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Models
- ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Thinking Augmented Pre-training
- A Multi-Agent Framework for Stateful Inference-Time Search
- You Need Better Attention Priors
查看其余 95 篇
- Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
- When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
- What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
- Vision Transformers are Circulant Attention Learners
- Understanding the Role of Training Data in Test-Time Scaling
- Understanding R1-Zero-Like Training: A Critical Perspective
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Towards Flash Thinking via Decoupled Advantage Policy Optimization
- Towards a Unified View of Large Language Model Post-Training
- Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction
- Think Outside the Policy: In-Context Steered Policy Optimization
- The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
- SpecAttn: Speculating Sparse Attention
- Sparse Attention Post-Training for Mechanistic Interpretability
- SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- Soft Adaptive Policy Optimization
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Sigmoid Loss for Language Image Pre-Training
- Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
- Sampling and Loss Weights in Multi-Domain Training
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
- Pre-training under infinite compute
- Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
- Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
- Optimizing Mixture of Block Attention
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- On the Convergence Rate of LoRA Gradient Descent
- Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning
- MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
- Modular Prompt Optimization: Optimizing Structured Prompts with Section-Local Textual Gradients
- Mixture-of-Depths Attention
- Mid-Training of Large Language Models: A Survey
- Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
- Memorization Dynamics in Knowledge Distillation for Language Models
- LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Language Self-Play For Data-Free Training
- KTO: Model Alignment as Prospect Theoretic Optimization
- Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
- Jailbroken: How Does LLM Safety Training Fail?
- Inpainting-Guided Policy Optimization for Diffusion Large Language Models
- In-Context Distillation with Self-Consistency Cascades: A Simple, Training-Free Way to Reduce LLM Agent Costs
- Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
- How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
- Higher-order Linear Attention
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
- Fast attention mechanisms: a tale of parallelism
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- End-to-End Test-Time Training for Long Context
- Efficient Streaming Language Models with Attention Sinks
- Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
- DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Defeating the Training-Inference Mismatch via FP16
- Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval
- Controlling changes to attention logits
- Continual Learning via Sparse Memory Finetuning
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- Better World Models Can Lead to Better Post-Training Performance
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training
- Accelerate Speculative Decoding with Sparse Computation in Verification
- A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving
- A Systematic Study of Model Merging Techniques in Large Language Models
- A Survey on LLM Mid-training
- A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
多模态与视觉 37 篇
- π_0: A Vision-Language-Action Flow Model for General Robot Control
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Visual Language Hypothesis
- Vision Transformers are Circulant Attention Learners
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Top 10 Open Challenges Steering the Future of Diffusion Language Model and Its Variants
- SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
- Sigmoid Loss for Language Image Pre-Training
- Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
查看其余 25 篇
- RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
- RewardDance: Reward Scaling in Visual Generation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- PaLM-E: An Embodied Multimodal Language Model
- OpenVLA: An Open-Source Vision-Language-Action Model
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Multimodal Deep Learning
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Mixture of Contexts for Long Video Generation
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- LinMU: Multimodal Understanding Made Linear
- Kimi K2.5: Visual Agentic Intelligence
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Inpainting-Guided Policy Optimization for Diffusion Large Language Models
- Improved Baselines with Visual Instruction Tuning
- Diffusion Language Models are Super Data Learners
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
- Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval
- BaseReward: A Strong Baseline for Multimodal Reward Model
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- A Survey on Multimodal Large Language Models
- A Survey on Agentic Multimodal Large Language Models
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
具身智能与机器人 10 篇
- π_0: A Vision-Language-Action Flow Model for General Robot Control
- DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- PaLM-E: An Embodied Multimodal Language Model
- OpenVLA: An Open-Source Vision-Language-Action Model
- Octo: An Open-Source Generalist Robot Policy
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- BridgeData V2: A Dataset for Robot Learning at Scale
- A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms
- A Comprehensive Survey on World Models for Embodied AI
AI安全与评测 29 篇
- PPTArena: A Benchmark for Agentic PowerPoint Editing
- A Unified Definition of Hallucination, Or: It's the World Model, Stupid
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Spanish Pre-trained BERT Model and Evaluation Data
- RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
- ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- OpenAssistant Conversations -- Democratizing Large Language Model Alignment
- LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- KTO: Model Alignment as Prospect Theoretic Optimization
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
查看其余 17 篇
- Jailbroken: How Does LLM Safety Training Fail?
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
- GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Extracting alignment data in open models
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- CreativityPrism: A Holistic Benchmark for Large Language Model Creativity
- Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
- BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
- AI Agent Systems: Architectures, Applications, and Evaluation
- A Survey on LLM-as-a-Judge
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- A Survey on Evaluation of Large Language Models
数据与AI工程 52 篇
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- Why Less is More (Sometimes): A Theory of Data Curation
- What's the next frontier for Data-centric AI? Data Savvy Agents
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- Understanding the Role of Training Data in Test-Time Scaling
- Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
- Train on Validation (ToV): Fast data selection with applications to fine-tuning
- Towards Automated Kernel Generation in the Era of LLMs
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
查看其余 40 篇
- SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
- Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
- Spanish Pre-trained BERT Model and Evaluation Data
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
- Open Data Synthesis For Deep Research
- Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
- Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
- Learning from Synthetic Data: Limitations of ERM
- Latent Traits and Cross-Task Transfer: Deconstructing Dataset Interactions in LLM Fine-tuning
- Language Self-Play For Data-Free Training
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
- Generative Data Refinement: Just Ask for Better Data
- Extracting alignment data in open models
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Educational data mining and learning analytics: An updated survey
- Diffusion Language Models are Super Data Learners
- Detecting Data Contamination in LLMs via In-Context Learning
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Dataset Growth
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Data-Efficient RLVR via Off-Policy Influence Guidance
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
- ConvergeWriter: Data-Driven Bottom-Up Article Construction
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- BridgeData V2: A Dataset for Robot Learning at Scale
- An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence
- Agentic Software Engineering: Foundational Pillars and a Research Roadmap
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- A Novel Combined Data-Driven Approach for Electricity Theft Detection
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
行业应用 21 篇
- LLM-ERM: Sample-Efficient Program Learning via LLM-Guided Search
- Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
- AI4X Roadmap: Artificial Intelligence for the advancement of scientific pursuit and its future directions
- Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
- Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction
- Students' Voices on Generative AI: Perceptions, Benefits, and Challenges in Higher Education
- Search over Self-Edit Strategies for LLM Adaptation
- Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
- Re4: Scientific Computing Agent with Rewriting, Resolution, Review and Revision
- QAgent: A modular Search Agent with Interactive Query Understanding
- On-line Policy Improvement using Monte-Carlo Search
- On GRPO Collapse in Search-R1: The Lazy Likelihood-Displacement Death Spiral
查看其余 9 篇
- MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
- LORE: A Large Generative Model for Search Relevance
- LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems
- Limits of trust in medical AI
- Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- Capabilities of GPT-4 on Medical Challenge Problems
- BloombergGPT: A Large Language Model for Finance
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
基础模型与理论 184 篇
- Step-GUI Technical Report
- Qwen3-VL Technical Report
- Monadic Context Engineering
- LLMs Encode How Difficult Problems Are
- LIMI: Less is More for Agency
- Large Language Models Meet Virtual Cell: A Survey
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Group Representational Position Encoding
- Fourier Neural Operators Explained: A Practical Perspective
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- An Augmentation Overlap Theory of Contrastive Learning
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
查看其余 172 篇
- xLLM Technical Report
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct
- Who Said Neural Networks Aren't Linear?
- What Affects the Effective Depth of Large Language Models?
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- Web World Models
- Virtual Width Networks
- VibeVoice Technical Report
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Understanding Robustness of Model Editing in Code LLMs: An Empirical Study
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Transition Models: Rethinking the Generative Learning Objective
- Transformers learn factored representations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Towards Unbiased Calibration using Meta-Regularization
- Towards Execution-Grounded Automated AI Research
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Tongyi DeepResearch Technical Report
- Thought Communication in Multiagent Collaboration
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
- The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- The Llama 3 Herd of Models
- The human biological advantage over AI
- The 2025 Foundation Model Transparency Index
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Structured Hints for Sample-Efficient Lean Theorem Proving
- Stronger Normalization-Free Transformers
- Step-DeepResearch Technical Report
- StarCoder: may the source be with you!
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- Sigmoid Head for Quality Estimation under Language Ambiguity
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Sentence-Anchored Gist Compression for Long-Context LLMs
- Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience
- Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Scaling and context steer LLMs along the same computational path as the human brain
- SAM 2: Segment Anything in Images and Videos
- s1: Simple test-time scaling
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
- Rethinking Cross-lingual Gaps from a Statistical Viewpoint
- Remote Labor Index: Measuring AI Automation of Remote Work
- Relative Scaling Laws for LLMs
- Relative-Based Scaling Law for Neural Language Models
- Reflect before Act: Proactive Error Correction in Language Models
- Recursive Language Models
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- Qwen2 Technical Report
- Quantitative Bounds for Length Generalization in Transformers
- QLoRA: Efficient Finetuning of Quantized LLMs
- Predicting Task Performance with Context-aware Scaling Laws
- PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- Opal: An Operator Algebra View of RLHF
- On the Origin of Algorithmic Progress in AI
- On the Fundamental Limits of LLMs at Scale
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- Object Recognition Datasets and Challenges: A Review
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- NRGPT: An Energy-based Alternative for GPT
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Nested Learning: The Illusion of Deep Learning Architectures
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
- Monitoring Monitorability
- Modeling Language as a Sequence of Thoughts
- Model Compression using Progressive Channel Pruning
- MobileLLM-Pro Technical Report
- Mixtral of Experts
- Mistral 7B
- MiMo-V2-Flash Technical Report
- Midtraining Bridges Pretraining and Posttraining Distributions
- mHC: Manifold-Constrained Hyper-Connections
- Mechanisms of Introspective Awareness
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- LLM Router: Prefill is All You Need
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Let's Verify Step by Step
- Learning to Discover at Test Time
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
- Large language models are not about language
- Large language models and the entropy of English
- Large Language Model Sourcing: A Survey
- Language models as tools for investigating the distinction between possible and impossible natural languages
- Kling-Omni Technical Report
- KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
- KAN: Kolmogorov-Arnold Networks
- Jailbreaking Black Box Large Language Models in Twenty Queries
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
- Introduction to Machine Learning
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Increasing the Thinking Budget is Not All You Need
- Improving Recursive Transformers with Mixture of LoRAs
- Improving Online Algorithms via ML Predictions
- HunyuanVideo 1.5 Technical Report
- Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- GPT-4o System Card
- GPT-4 Technical Report
- Geometric and Dynamic Scaling in Deep Transformers
- Generative Early Stage Ranking
- Generative AI
- Gemma 2: Improving Open Language Models at a Practical Size
- FLEx: Language Modeling with Few-shot Language Explanations
- F -- A Model of Events based on the Foundational Ontology DOLCE+DnS Ultralite
- Explaining the Success of Nearest Neighbor Methods in Prediction
- Expand Neurons, Not Parameters
- Excess Description Length of Learning Generalizable Predictors
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Epistemological Fault Lines Between Human and Artificial Intelligence
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Emergent Introspective Awareness in Large Language Models
- ELLA: Efficient Lifelong Learning for Adapters in Large Language Models
- Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models
- DELTA: Decoupling Long-Tailed Online Continual Learning
- DeepSeek-V3 Technical Report
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Deep sequence models tend to memorize geometrically; it is unclear why
- Deep Delta Learning
- Context-Free Recognition with Transformers
- Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
- CaveAgent: Transforming LLMs into Stateful Runtime Operators
- Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
- Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
- Bias and Fairness in Large Language Models: A Survey
- Beyond the Black Box: Theory and Mechanism of Large Language Models
- Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale
- Behind RoPE: How Does Causal Mask Encode Positional Information?
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- BEAM: Brainwave Empathy Assessment Model for Early Childhood
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Are Large Language Models Sensitive to the Motives Behind Communication?
- An Information-Theoretic Framework for Robust Large Language Model Editing
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- Allocation of Parameters in Transformers
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Algorithmic Thinking Theory
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- Action Language BC+
- Accurate Table Question Answering with Accessible LLMs
- A Survey of Weight Space Learning: Understanding, Representation, and Generation
- A Survey of Vibe Coding with Large Language Models
- A Survey of Large Language Models
- A Survey of AI Scientists: Surveying the automatic Scientists and Research
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- A model of errors in transformers
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- A Definition of AGI
- A Concise Review of Hallucinations in LLMs and their Mitigation
- A Component-Based Survey of Interactions between Large Language Models and Multi-Armed Bandits