Skip to content
Home/ Reinforcement Learning and Stochastic Optimization: A Unified Framework for Sequential Decisions
Reinforcement Learning and Stochastic Optimization: A Unified Framework for Sequential Decisions

Reinforcement Learning and Stochastic Optimization: A Unified Framework for Sequential Decisions

No customer reviews yet ISBN 9781119815037 Wiley

REINFORCEMENT LEARNING AND STOCHASTIC OPTIMIZATION

Clearing the jungle of stochastic optimization

Sequential decision problems, which consist of "decision, information, decision, information," are ubiquitous, spanning virtually every human activity ranging from business applications, health (personal and public health, and medical decision making), energy, the sciences, all fields of engineering, finance, and e-commerce. The diversity of applications attracted the attention of at least 15 distinct fields of research, using eight distinct notational systems which produced a vast array of analytical tools. A byproduct is that powerful tools developed in one community may be unknown to other communities.

Reinforcement Learning and Stochastic Optimization offers a single canonical framework that can model any sequential decision problem using five core components: state variables, decision variables, exogenous information variables, transition function, and objective function. This book highlights twelve types of uncertainty that might enter any model and pulls together the diverse set of methods for making decisions, known as policies, into four fundamental classes that span every method suggested in the academic literature or used in practice.

Reinforcement Learning and Stochastic Optimization is the first book to provide a balanced treatment of the different methods for modeling and solving sequential decision problems, following the style used by most books on machine learning, optimization, and simulation. The presentation is designed for readers with a course in probability and statistics, and an interest in modeling and applications. Linear programming is occasionally used for specific problem classes. The book is designed for readers who are new to the field, as well as those with some background in optimization under uncertainty.

Throughout this book, readers will find references to over 100 different applications, spanning pure learning problems, dynamic resource allocation problems, general state-dependent problems, and hybrid learning/resource allocation problems such as those that arose in the COVID pandemic. There are 370 exercises, organized into seven groups, ranging from review questions, modeling, computation, problem solving, theory, programming exercises and a "diary problem" that a reader chooses at the beginning of the book, and which is used as a basis for questions throughout the rest of the book.

About the author

Product details

BrandWiley
Pub dateMar 15, 2022
ISBN-101119815037
ISBN-139781119815037
Hardcover1136.0 pages
LanguageEnglish
Dimensions0.39 × 0.39 × 0.39 in
Weight4 lb
Last updated 2026-08-06 19:12
$140.59 $161.95 13% off
You save $21.36 · list price $161.95
In stock soon — order now to reserve your copy
Delivery by Monday, September 14, 2026
Qty
Sign in to Add to Saved list
Free delivery on orders over $35.
15-day returns. Any reason.
Secure checkout. We never store card details.

Readers who bought this also bought

More from Linear & Nonlinear Programming
See all
Multicriteria Optimization in Engineering and in the Sciences (1988)
We are rarely asked to. make decisions based on only one criterion&#x3b; most often, decisions are based on several usually confticting, criteria. In nature, if the design of a system evolves to some final, optimal state, then it must include a balance for the interaction of the system with its surroundings- certainly a design based on a variety of criteria. Furthermore, the diversity of nature's designs suggests an infinity of such optimal states. In another sense, decisions simultaneously optimize a finite number of criteria, while there is usually an infinity of optimal solutions. Multicriteria optimization provides the mathematical framework to accommodate these demands. Multicriteria optimization has its roots in mathematical economics, in particular, in consumer economics as considered by Edgeworth and Pareto. The critical question in an exchange economy concerns the "equilibrium point" at which each of N consumers has achieved the best possible deal for hirnself or herself. Ultimately, this is a collective decision in which any further gain by one consumer can occur only at the expense of at least one other consumer. Such an equilibrium concept was first introduced by Edgeworth in 1881 in his book on mathematical psychics. Today, such an optimum is variously called "Pareto optimum" (after the Italian-French welfare economist who continued and expanded Edgeworth's work), "effi. cient," "nondominated," and so on.
$177.01