About

I am a Principal Applied Scientist at Amazon, where I lead research on large-scale reinforcement learning for supply-chain decision making. My work focuses on learning systems that coordinate large populations of agents, operate under shared constraints, and make good decisions in complex real-world environments.

More recently, I have been applying that perspective to language models and autonomous agents, with a focus on how models reason, improve through reinforcement learning, and act reliably in increasingly open-ended settings.

I received my PhD from Princeton University in 2019, advised by Han Liu.

Language models and agents

  • Reasoning and self-improvement. When can LLMs improve themselves? A formalization through the generation-verification gap, and how it scales with pretraining compute. Separately, structured intermediate representations, from functional programs to formal specifications, that make models markedly better at generating verified code and proofs, both at inference time and as post-training targets. [12]
  • RL post-training. Learned decoding policies that allocate test-time compute over a model's token tree, trained with behavior cloning and RL, and curriculum design for RL fine-tuning (in progress).
  • Agents. L3M, a multi-agent runtime and coding-agent harness written in Lean 4 whose safety boundaries are machine-checked theorems (open-source release and white paper coming soon), and measuring and improving agentic systems end to end. [12]

Reinforcement learning and deep learning

  • Multi-agent coordination and capacity control. Learned dual prices coordinating hundreds of thousands of agents; deep RL for inventory management in production. [1234]
  • Policy gradients. Variance-reduced estimators for structured action spaces; natural policy gradient for exponential families. [12]
  • Differentiable optimizers and sim-to-real. Value functions inside convex optimization layers; differentiable simulators and closing the sim-to-real gap. [123]
  • Forecasting and statistical learning. Multi-horizon probabilistic forecasting; certifiably optimal clustering and latent-variable graphical models. [1234]

News

Publications and preprints

  1. L3M: Capability Without Authority. A Verifiable Harness and Runtime for Self-Extending and Self-Improving Agents
    Carson Eisenach, Robert Joseph George, Vincent Quenneville-Bélair, Dean Foster
    White paper (release forthcoming), 2026
    Abstract

    The harness that turns a language model's output into file edits, processes, and network calls holds all of the authority the model appears to exercise, yet in most systems it is an afterthought: a permission dialog and a hope that the model behaves. L3M bounds that authority by proof. The harness is a Lean 4 program, so its safety properties are kernel-checked theorems about the code that runs: file access is confined to an agent's roots, a child's grant never exceeds its parent's, budgets are lent rather than minted, and irreversible actions are always reviewed. Agents extend their own tool surface with sealed programs whose types state and bound what they may do, so capability grows while authority does not. Reversibility, derived from each tool's footprint, decides what an agent may do without asking, and review is conducted as deliberation between agents with the human kept in the loop. A Git-backed distributed runtime makes the permission, ref, and directory trees one tree, so guarantees proved for an agent hold for its whole subtree, and agent behaviour is a fold over the log, which lets experiments that measure and improve agents run inside the harness as ordinary bounded agents. The formalization is incremental: every headline guarantee is stated in readable terms and tagged with what it rests on, computed from its proof.

  2. Harrison Waldon*, Carson Eisenach*, Akhil Bagaria, Daniel Russo, Dominique Perrault-Joncas, Alisha Zachariah, Dean Foster
    arXiv preprint, 2026arxiv
    Abstract

    We study two-timescale decision systems in which a planning layer periodically supplies a continuation-value function to a real-time optimizer that allocates arriving resources, with inventory placement as our motivating application. We propose an end-to-end reinforcement learning (RL) method for learning this function using proximal residual value functions, which combine a strictly convex potential of post-decision inventory with a learned convex residual. This general form yields a well-posed optimization layer that supports end-to-end differentiation while preserving an explicit convex objective for real-time execution. We characterize the necessary and sufficient conditions under which a smooth value function yields decisions that are consistent across the planning and execution timescales. In an offline simulation using historical inventory arrival and demand patterns from a large e-commerce retailer, learned proximal residual value functions reduce total routing and transfer cost relative to a historical-production-system proxy by 5.0%.

  3. Kate Gwimm, Carson Eisenach
    COLM 2026 Workshop on Context Beyond the Windowarxiv
    Abstract

    Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands of tables, and no model can ingest a full catalog at once. We argue that the most effective place to intervene is therefore the knowledge-base context the model consumes, and that this context should be constructed from historical usage rather than tuned for as a fixed input. Using a query-DAG decomposition recovered from production SQL, we compare the value of oracle query graphs versus retrieved knowledge-base context, and optimize a distillation procedure that turns historical query profiles into reusable SQL reference cards. On a benchmark of 5176 production queries from a major online retailer, optimizing these context artifacts yields larger gains (~12-25% AST similarity) than optimizing the retrieval harness (~3-12%).

  4. Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar, Dean Foster, Carson Eisenach
    arXiv preprint, 2026arxiv
    Abstract

    In large-scale multi-agent systems with shared resource constraints, an upstream planner must iteratively evaluate candidate resource plans before committing to one. Lagrangian relaxation separates local decisions through a broadcast cost signal, but the planner still needs the cost-to-utilization response map, which depends on population composition that changes across planning cycles. We propose population-aware coordination interfaces: learned primal and dual maps, conditioned on compact population summaries, that the planner queries inside its iterative loop. These maps remain reliable across evolving populations without per-cycle retraining, and support coordination of large populations from compact subsamples. In a supply-chain capacity-control case study, population-aware interfaces reduce forecast error by 16-19% and capacity violations by 20-51% relative to population-unaware baselines under composition shift; 20K-agent cohorts support accurate coordination of 500K-agent populations.

  5. Robert Joseph George, Carson Eisenach, Udaya Ghai, Dominique Perrault-Joncas, Anima Anandkumar, Dean Foster
    ICML 2026 AI4Math Workshoparxiv
    Abstract

    BRIDGE is a framework for improving verified program synthesis with large language models. The approach decomposes verification into three interconnected domains: Code (implementations), Specifications (formal intent), and Theorem Statements. BRIDGE uses a code-first workflow, using the generated implementation as a semantic anchor for downstream specification and theorem statement generation. Evaluated across five LLMs on 178 algorithmic problems, the method achieves nearly 1.5x better Lean executable correctness compared to baselines and demonstrates 2x more sample-efficient inference. The framework also improves Python pass rates by up to 17.5% through specification-driven prompting. Supervised fine-tuning with BRIDGE-style reasoning yields nearly 1.5x higher Lean pass success than code-only SFT.

  6. Alvaro Maggiar, Sohrab Andaz, Akhil Bagaria, Carson Eisenach, Dean Foster, Omer Gottesman, Dominique Perrault-Joncas
    NeurIPS 2025 MLxOR Workshoparxiv
    Abstract

    This paper investigates the application of Deep Reinforcement Learning (DRL) to classical inventory management problems, with a focus on practical implementation considerations. We apply a DRL algorithm based on DirectBackprop to several fundamental inventory management scenarios including multi-period systems with lost sales (with and without lead times), perishable inventory management, dual sourcing, and joint inventory procurement and removal. The DRL approach learns policies across products using only historical information that would be available in practice, avoiding unrealistic assumptions about demand distributions or access to distribution parameters. We demonstrate that our generic DRL implementation performs competitively against or outperforms established benchmarks and heuristics across these diverse settings, while requiring minimal parameter tuning. Through examination of the learned policies, we show that the DRL approach naturally captures many known structural properties of optimal policies derived from traditional operations research methods. To further improve policy performance and interpretability, we propose a Structure-Informed Policy Network technique that explicitly incorporates analytically-derived characteristics of optimal policies into the learning process. This approach can help interpretability and add robustness to the policy in out-of-sample performance, as we demonstrate in an example with realistic demand data. Finally, we provide an illustrative application of DRL in a non-stationary setting. Our work bridges the gap between data-driven learning and analytical insights in inventory management while maintaining practical applicability.

  7. Riccardo Savorgnan, Udaya Ghai, Carson Eisenach, Dean Foster
    arXiv preprint, 2025arxiv
    Abstract

    This paper addresses forecasting the volume of inventory units fulfilled from each warehouse and associated shipping costs. We model the joint distribution of outbound drain and costs across warehouses, conditioned on inventory positions and customer demand. A key challenge is ensuring the model is differentiable for use within reinforcement learning simulators, since actual production systems are too slow and non-differentiable for RL training. We propose a validation scheme leveraging production systems to evaluate model performance on counterfactual inventory states generated by RL policies, demonstrating accuracy in in-distribution settings.

  8. Yuda Song, Hanlin Zhang, Carson Eisenach, Sham Kakade, Dean Foster, Udaya Ghai
    International Conference on Learning Representations (ICLR 2025)arxiv
    Abstract

    Self-improvement is a mechanism in Large Language Model (LLM) pre-training, post-training and test-time inference. We explore a framework where the model verifies its own outputs, filters or reweights data based on this verification, and distills the filtered data. Despite several empirical successes, a fundamental understanding is still lacking. In this work, we initiate a comprehensive, modular and controlled study on LLM self-improvement. We provide a mathematical formulation for self-improvement, which is largely governed by a quantity which we formalize as the generation-verification gap. Through experiments with various model families and tasks, we discover a scaling phenomenon of self-improvement -- a variant of the generation-verification gap scales monotonically with the model pre-training flops. We also examine when self-improvement is possible, an iterative self-improvement procedure, and ways to improve its performance. Our findings not only advance understanding of LLM self-improvement with practical implications, but also open numerous avenues for future research into its capabilities and boundaries.

  9. Carson Eisenach, Udaya Ghai, Dhruv Madeka, Kari Torkkola, Dean Foster, Sham Kakade
    arXiv preprint, 2024arxiv
    Abstract

    This paper addresses the capacitated periodic review inventory control problem, focusing on a retailer managing multiple products with limited shared resources, such as storage or inbound labor at a facility. Specifically, this paper is motivated by the questions of (1) what does it mean to backtest a capacity control mechanism, (2) can we devise and backtest a capacity control mechanism that is compatible with recent advances in deep reinforcement learning for inventory management? First, because we only have a single historic sample path of Amazon's capacity limits, we propose a method that samples from a distribution of possible constraint paths covering a space of real-world scenarios. This novel approach allows for more robust and realistic testing of inventory management strategies. Second, we extend the exo-IDP (Exogenous Decision Process) formulation of Madeka et al. 2022 to capacitated periodic review inventory control problems and show that certain capacitated control problems are no harder than supervised learning. Third, we introduce a `neural coordinator', designed to produce forecasts of capacity prices, guiding the system to adhere to target constraints in place of a traditional model predictive controller. Finally, we apply a modified DirectBackprop algorithm for learning a deep RL buying policy and a training the neural coordinator. Our methodology is evaluated through large-scale backtests, demonstrating RL buying policies with a neural coordinator outperforms classic baselines both in terms of cumulative discounted reward and capacity adherence (we see improvements of up to 50% in some cases).

  10. Sohrab Andaz, Carson Eisenach, Dhruv Madeka, Kari Torkkola, Randy Jia, Dean Foster, Sham Kakade
    GenAI4DM Workshop, ICLR 2024arxiv
    Abstract

    In this paper we address the problem of learning and backtesting inventory control policies in the presence of general arrival dynamics -- which we term as a quantity-over-time arrivals model (QOT). We also allow for order quantities to be modified as a post-processing step to meet vendor constraints such as order minimum and batch size constraints -- a common practice in real supply chains. To the best of our knowledge this is the first work to handle either arbitrary arrival dynamics or an arbitrary downstream post-processing of order quantities. Building upon recent work (Madeka et al., 2022) we similarly formulate the periodic review inventory control problem as an exogenous decision process, where most of the state is outside the control of the agent. Madeka et al. (2022) show how to construct a simulator that replays historic data to solve this class of problem. In our case, we incorporate a deep generative model for the arrivals process as part of the history replay. By formulating the problem as an exogenous decision process, we can apply results from Madeka et al. (2022) to obtain a reduction to supervised learning. Finally, we show via simulation studies that this approach yields statistically significant improvements in profitability over production baselines. Using data from an ongoing real-world A/B test, we show that Gen-QOT generalizes well to off-policy data.

  11. Dhruv Madeka, Kari Torkkola, Carson Eisenach, Anna Luo, Dean Foster, Sham Kakade
    arXiv preprint, 2022arxiv
    Abstract

    This work provides a Deep Reinforcement Learning approach to solving a periodic review inventory control system with stochastic vendor lead times, lost sales, correlated demand, and price matching. While this dynamic program has historically been considered intractable, our results show that several policy learning approaches are competitive with or outperform classical methods. In order to train these algorithms, we develop novel techniques to convert historical data into a simulator. On the theoretical side, we present learnability results on a subclass of inventory control problems, where we provide a provable reduction of the reinforcement learning problem to that of supervised learning. On the algorithmic side, we present a model-based reinforcement learning procedure (Direct Backprop) to solve the periodic review inventory control problem by constructing a differentiable simulator. Under a variety of metrics Direct Backprop outperforms model-free RL and newsvendor baselines, in both simulations and real-world deployments.

  12. Sitan Yang, Carson Eisenach, Dhruv Madeka
    KDD MiLeTS Workshop 2022pdfarxiv
    Abstract

    Multi-horizon probabilistic time series forecasting has wide applicability to real-world tasks such as demand forecasting. Recent work in neural time-series forecasting mainly focus on the use of Seq2Seq architectures. For example, MQTransformer - an improvement of MQCNN - has shown the state-of-the-art performance in probabilistic demand forecasting. In this paper, we consider incorporating cross-entity information to enhance model performance by adding a cross-entity attention mechanism along with a retrieval mechanism to select which entities to attend over. We demonstrate how our new neural architecture, MQRetNN, leverages the encoded contexts from a pretrained baseline model on the entire population to improve forecasting accuracy. Using MQCNN as the baseline model (due to computational constraints, we do not use MQTransformer), we first show on a small demand forecasting dataset that it is possible to achieve ~3% improvement in test loss by adding a cross-entity attention mechanism where each entity attends to all others in the population. We then evaluate the model with our proposed retrieval methods - as a means of approximating an attention over a large population - on a large-scale demand forecasting application with over 2 million products and observe ~1% performance gain over the MQCNN baseline.

  13. Carson Eisenach, Florentina Bunea, Yang Ning, Claudiu Dinicu
    Journal of Machine Learning Research, 21 (2020)pdfarxiv
    Abstract

    Motivated by modern applications in which one constructs graphical models based on a very large number of features, this paper introduces a new class of cluster-based graphical models. Unlike standard graphical models, variable clustering is applied as an initial step for reducing the dimension of the feature space. We employ model assisted clustering, in which the clusters contain features that are similar to the same unobserved latent variable. Two different cluster-based Gaussian graphical models are considered: the latent variable graph, corresponding to the graphical model associated with the unobserved latent variables, and the cluster-average graph, corresponding to the vector of features averaged over clusters. We derive estimates tailored to these graphs, with the goal of pattern recovery under false discovery rate (FDR) control. Our study reveals that likelihood based inference for the latent graph is analytically intractable, and we develop alternative estimation and inference strategies. We replace the likelihood of the data by appropriate empirical risk functions that allow for valid inference in both graphical models under study. Our main results are Berry-Esseen central limit theorems for the proposed estimators, which are proved under weaker assumptions than those employed in the existing literature on Gaussian graphical model inference. We make explicit the implications of the asymptotic approximations on graph recovery under FDR control, and show when it can be controlled asymptotically. Our analysis takes into account the uncertainty induced by the initial clustering step. We find that the errors induced by clustering are asymptotically ignorable in the follow-up analysis, under no further restrictions on the parameter space for which inference is valid. The theoretical properties of the proposed procedures are verified on simulated data and an fMRI data analysis.

  14. Carson Eisenach, Yagna Patel, Dhruv Madeka
    arXiv preprint, 2020arxivicml workshopkdd workshop pdfAlso presented at the ICML 2022 Workshop on Continuous Time Perspectives in Machine Learning and at the KDD 2022 MiLeTS Workshop, as "MQTransformer: Multi-Horizon Forecasts with Context Dependent Attention and Optimal Bregman Volatility" (with Kevin Chen and Lee Dicker).
    Abstract

    Recent advances in neural forecasting have produced major improvements in accuracy for probabilistic demand prediction. In this work, we propose novel improvements to the current state of the art by incorporating changes inspired by recent advances in Transformer architectures for Natural Language Processing. We develop a novel decoder-encoder attention for context-alignment, improving forecasting accuracy by allowing the network to study its own history based on the context for which it is producing a forecast. We also present a novel positional encoding that allows the neural network to learn context-dependent seasonality functions as well as arbitrary holiday distances. Finally we show that the current state of the art MQ-Forecaster (Wen et al., 2017) models display excess variability by failing to leverage previous errors in the forecast to improve accuracy. We propose a novel decoder-self attention scheme for forecasting that produces significant improvements in the excess variation of the forecast.

  15. Carson Eisenach, Han Liu
    Mathematical Programming Series B (2019)publisherarxiv
    Abstract

    We consider SDP relaxation methods for data and variable clustering problems, which have been shown in the literature to have good statistical properties in a variety of settings, but remain intractable to solve in practice. In particular, we propose FORCE, a new algorithm to solve the Peng-Wei $K$-means SDP. Compared to the naive interior point method, our method reduces the computational complexity of solving the SDP from $\\tilde{O}(d^7 \\log \\epsilon^{-1})$ to $\\tilde{O}(d^{6}K^{-2}\\epsilon^{-1})$. Our method combines a primal first-order method with a dual optimality certificate search, which when successful, allows for early termination of the primal method. We show under certain data generating distributions that, with high probability, FORCE is guaranteed to find the optimal solution to the SDP relaxation and provide a certificate of exact optimality. As verified by our numerical experiments, this allows FORCE to solve the Peng-Wei SDP with dimensions in the hundreds in only tens of seconds. We also consider a variation of the Peng-Wei SDP for the case when $K$ is not known a priori and show that a slight modification of FORCE reduces the computational complexity of solving this problem as well: from $\\tilde{O}(d^7 \\log \\epsilon^{-1})$ using a standard SDP solver to $\\tilde{O}(d^{4}\\epsilon^{-1})$.

  16. Carson Eisenach, Haichuan Yang, Ji Liu, Han Liu
    International Conference on Learning Representations (ICLR 2019)arxivcode
    Abstract

    Many complex domains, such as robotics control and real-time strategy (RTS) games, require an agent to learn a continuous control. In the former, an agent learns a policy over $\mathbb{R}^d$ and in the latter, over a discrete set of actions each of which is parametrized by a continuous parameter. Such problems are naturally solved using policy based reinforcement learning (RL) methods, but unfortunately these often suffer from high variance leading to instability and slow convergence. We show that in many cases a substantial portion of the variance in policy gradient estimators is completely unnecessary and can be eliminated without introducing bias. Unnecessary variance is introduced whenever policies over bounded action spaces are modeled using distributions with unbounded support, by applying a transformation T to the sampled action before execution in the environment. Recent works have studied variance reduced policy gradients for actions in bounded intervals, but to date no variance reduced methods exist when the action is a direction -- constrained to the unit sphere -- something often seen in RTS games. To address these challenges we: (1) introduce a stochastic policy gradient method for directional control; (2) introduce the marginal policy gradient framework, a powerful technique to obtain variance reduced policy gradients for arbitrary $T$; (3) show that marginal policy gradients are guaranteed to reduce variance, quantifying that reduction exactly; (4) validate our framework by applying the methods to a popular RTS game and a navigation task, demonstrating improvement over a policy gradient baseline.

  17. Carson Eisenach, Zhuoran Yang
    Technical report, 2018pdfcode
    Abstract

    Recent work has highlighted how a misalignment between the support of the policy and the action space of the reinforcement learning problem can introduce bias and unnecessary variance into policy gradient estimates. To better align the support of the policy and the action space, we can consider using arbitrary exponential families to model the policy distribution. Exponential families are a natural choice because the class of exponential families is very rich and can model the support of most action spaces of practical interest. While the multivariate Gaussian is the most commonly used distribution today, in general it is possible to efficiently implement both natural policy gradient and TRPO for any exponential family. In this technical report we derive efficient natural policy gradient update rules for several exponential familes. We also consider an application of the Gamma distribution to an optimal production problem and show that it substantially outperforms the Gaussian.

  18. Carson Eisenach, Zhaoran Wang, Han Liu
    Technical report, 2017pdf
    Abstract

    We provide a principled framework for non-parametrically learning activation functions in deep neural networks. Currently, state-of-the-art deep networks treat choice of activation function as a hyper-parameter before training. By allowing activation functions to be estimated as part of the training procedure, we expand the class of functions that each node in the network can learn. Building on recent advances in stability bounds for stochastic gradient methods, we provide a theoretical justification for our choice of nonparametric activation functions and demonstrate that networks with our nonparametric activation functions generalize well. To demonstrate the power of our novel techniques, we test them on image recognition datasets and achieve up to a 15% relative increase in test performance compared to the baseline.

Education