IEE 475: Simulating Stochastic Systems: Recent Episodes

Theodore P. Pavlic

Archived lectures from undergraduate course on stochastic simulation given at Arizona State University by Ted Pavlic

View Details

During this lecture, we review the topics covered up to this point in the course as preparation for the upcoming midterm exam. Students are encouraged to bring their own questions to class so that we can focus on the topics that students feel like they need the most help with.

View Details

In this lecture, we review pseudo-random number generation and then introduce random-variate generation by way of inverse-transform sampling. In particular, we start with a review of the two most important properties of a pseudo-random number generator (PRNG), uniformity and independence, and discuss statistically rigorous methods for testing for these two properties. For uniformity, we focus on a Chi-square/Chi-squared test for larger numbers of samples and a Kolmogorov–Smirnov (KS) test for smaller numbers of samples. For independence, we discuss autocorrelation tests and runs test, and then we demonstrate a runs above-and-below-the-mean test. We then shift to discussing inverse-transform sampling for continuous random variates and discrete random variates and how the resulting random-variate generators might be implemented in a tool like Rockwell Automation's Arena.

View Details

In this lecture, we first cover some discrete distributions (and the Poisson process) that we ran out of time for during the previous lecture. We then launch into a discussion of how to generate pseudo-random numbers distributed uniformly between 0 and 1 (which are necessary for us to easily generate random variates of any distribution). We talk about the two most important properties of a pseudo-random number generator (PRNG), uniformity and independence. We then talk about desirable properties. Some examples are given of some early PRNG's, and then we introduce the linear congruential generator (LCG) and its variants (like the Combined Linear Congruential Generator, CLCG), which represent a much more modern PRNG that has a number of good properties. We close with a discussion of tests of uniformity. We will continue this discussion and add on tests for independence during next lecture (which will primarily cover random-VARIATE generation).

View Details

In this lecture, we review basic probability fundamentals (measure spaces, probability measures, random variables, probability density functions, probability mass functions, cumulative distribution functions, moments, mean/expected value/center of mass, standard deviation, variance), and then we start to build a vocabulary of different probabilistic models that are used in different modeling contexts. These include uniform, triangular, normal, exponential, Erlang-k, Weibull, and Poisson variables. We will finish the discussing next time with the Bernoulli-based discrete variables and Poisson processes.

View Details

In this lecture, we introduce the measure-theoretic concept of a random variable (which is neither random nor a variable) and related terms, such as outcomes, events, probability measures, moments, means, etc. Throughout the lecture, we use the metaphor of probability as mass (and thus probability density as mass density, and a mean as a center of mass). This allows us to discuss the "statistical leverage" of outliers in a distribution (i.e., although they happen infrequently, they still have the ability to shift the mean significantly, as in physical leverage). This sets us up to talk about random processes and particular random variables in the next lecture.

View Details

This lecture provides some historical background and motivation for System Dynamics Modeling (SDM) and Agent-Based Modeling (ABM), two other simulation modeling approaches that contrast with Discrete Event System (DES) simulation.

In particular, in this lecture, we briefly introduce System Dynamics Modeling (SDM) and Agent-Based/Individual-Based Modeling (ABM/IBM) as the two ends of the simulation modeling spectrum (from low resolution to high resolution). The introduction of ABM describes applications in life sciences, social sciences, and engineering (Multi-Agent Systems, MAS)/operations research. NetLogo is introduced, and it is used to present examples of running ABM's as well as the code behind them.

This lecture is also be coupled with notes discussing the Lab 3 (Monte Carlo simulation) results and general experience. These comments focus on interval estimation (which is right 95% of the time, as opposed to point estimation that is right 0% of the time) and the role of non-trivial distributions of random variables (as opposed to just their means).

View Details

This lecture covers content related to implementing simulations with spreadsheets and the motivations for the use of special-purpose Discrete Event System Simulation tools. In particular, we discuss different approaches to implementing Discrete Event System (DES) simulations (DESS) with simple spreadsheets (e.g., Microsoft Excel, Google Sheets, Apple Numbers, etc.). We cover inventory management problems (such as the newsvendor model) as well as Monte Carlo sampling and stochastic activity networks (SAN's). Although we show that spreadsheets can be very powerful for this kind of work, we highlight that this approach is cumbersome for systems with increasing complexity. So this motivates why we would use more sophisticated tools specifically built for simulation (but perhaps not so great for data analysis by themselves), like Arena, FlexSim, Simio, and NetLogo.

View Details

In this lecture, we close out our review of DES fundamentals and hand simulation. After going through a hand-simulation example one last time, we show how to implement a Discrete Event System (DES) simulation using a spreadsheet tool like Microsoft Excel without any "macros" (VBA, etc.). This involves defining relationships ACROSS TIME that allow the spreadsheet to (in a declarative fashion) reconstruct the trajectory that is the output of the simulation.

At the end of the lecture, we pivot to discussing the previous "Lab 2 (Muffin Oven Simulation)", which lets us introduce common random numbers (CRNs), statistical blocking, requirements of 2-sample and paired t-tests, and more sophisticated statistical methods that better characterize PRACTICAL significance (and take into account the multiple comparisons problem). Thus, the post-lab2 reflections are largely a preview of future topics in the course.

View Details

In this lecture, we review fundamentals of Discrete Event System (DES) simulation (e.g., entities, resources, activities, processes, delays, attributes) and we run through a number of DES modeling examples. These examples show how different research/operations questions can lead to different choices of entities/resources/etc. We close with a hand-simulation example of a single-channel, single-server queue with provided interarrival times and service times.

View Details

In this lecture, we cover fundamentals of discrete-event system (DES) simulation (DESS). This involves reviewing basic simulation concepts (entities, resources, attributes, events, activities, delays) and introducing the event-scheduling world view, which provides a causality framework on which an automatic simulation of a DES system can be built. We also discuss briefly how the stochastic modeling inherent to DESS means that outputs will be variable and thus will require rigorous statistics to make sense of.

View Details

In this lecture, we introduce the three different simulation methodologies (agent-based modeling, system dynamics modeling, and discrete event system simulation) and then focus on how stochastic modeling is used within discrete-event system simulation. In particular, we define terms such as system, dynamic system, state, state variable, activity, delay, resource, entity, and the notion of "input modeling."

View Details

In this lecture, we introduce Industrial and Systems Engineering as a blend of science and engineering that necessitates model building. We then define model (as something that answers a "What If" question) and different types of models. This gives us an opportunity to discuss how modeling is less about describing reality and more about generating tools to do useful things/make useful predictions. We end with a comparison of mental and quantitative models, as well as a comparison of different types of quantitative models (including simulation modeling).

View Details

This lecture introduces students to IEE 475 (Simulating Stochastic Systems), a required course for Industrial Engineering majors that covers the design and analysis of simulation models of real-world engineered systems. The lecture covers contents of the syllabus as well as where students can find more information in the Canvas Learning Management System site for the course.

View Details

In this lecture, we prepare for the final exam and give a brief review of all topics from the course. Students are encouraged to bring their own questions so that the focus of the class is on the topics that students feel they need the most help with.

View Details

In this lecture, we wrap up the course content in IEE 475. We first do a quick overview of the four variance reduction techniques (VRT's) covered in Unit K. That is, we cover: common random numbers (CRN's), antithetic variates (AV's), importance sampling, and control variates. We then remember some general comments about the goal of modeling and commonalities seen across simulation platforms (as well as the different types of simulation platforms in general).

View Details

In this lecture, we review four different Variance Reduction Techniques (VRT's). Namely, we discuss common random numbers (CRNs), control variates, antithetic variates (AVs), and importance sampling. Each one of these is a different approach to reducing the variance in the estimation of relative or absolute performance of a simulation model. Variance reduction is an alternative way to increase the power of a simulation that is hopefully less costly than increasing the number of replications.

View Details

In this lecture, we start by reviewing approaches for absolute and relative performance estimation in stochastic simulation. This begins with a reminder of the use of confidence intervals for estimation of performance for a single simulation model. We then move to different ways to use confidence intervals on mean DIFFERENCES to compare two different simulation models. We then move to the ranking and selection problem for three or more different simulation models, which allows us to talk about analysis of variance (ANOVA) and post hoc tests (like the Tukey HSD or Fisher's LSD). After that review, we move on to introducing variance reduction techniques (VRTs) which reduce the size of confidence intervals by experimentally controlling/accounting for alternative sources of variance (and thus reducing the observed variance in response variables). We discuss Common Random Numbers (CRNs), which use a paired/blocked design to reduce the variance caused by different random-number streams. We start to discuss control variates (CVs), but that discussion will be picked up at the start of the next lecture.

View Details

In this lecture, we review what we have learned about one-sample confidence intervals (i.e., how to use them as graphical versions of one-sample t-tests) for absolute performance estimation in order to motivate the problem of relative performance estimation. We introduce two-sample confidence intervals (i.e., confidence intervals on DIFFERENCES based on different two-sample t-tests) that are tested against a null hypothesis of 0. This means covering confidence interval half widths for the paired-difference t-test, the equal-variance (pooled) t-test, and Welch's unequal variance t-test. Each of these different experimental conditions sets up a different standard error of the mean formula and formula for degrees of freedom that are used to define the actual confidence interval half widths (centered on the difference in sample means in the pairwise comparison of systems). We then generalize to the case of more than 2 systems, particularly for "ranking and selection (R&S)." This lets us review the multiple-comparisons problem (and Bonferroni correction) and how post hoc tests (after an ANOVA) are more statistically powerful ways to do comparisons.

View Details

In this lecture, we start by further reviewing confidence intervals (where they come from and what they mean) and prediction intervals and then use them to motivate a simpler way to determine how many replications are needed in a simulation study (focusing first on transient simulations of terminating systems). We then shift our attention to steady-state simulations of non-terminating systems and the issue of initialization bias. We discuss different methods of "warming up" a steady-state simulation to reduce initialization bias and then merge that discussion with the prior discussion on how to choose the number of replications. In the next lecture, we'll finish up with a discussion of the method of "batch means" in steady-state simulations.

View Details

In this lecture, we review estimating absolute performance from simulation, with focus on choosing the number of necessary replications of transient simulations of terminating systems. The lecture starts by overviewing point estimation, bias, and different types of point estimators. This includes an overview of quantile estimation and how to use quantile estimation to use simulations as null-hypothesis-prediction generators. We the introduce interval estimation with confidence intervals and prediction intervals. Confidence intervals, which are visualizations of t-tests, provide an alternative way to choose the number of required replications without doing a formal power analysis.

View Details

In this lecture, we introduce the estimation of absolute performance measures in simulation – effectively shifting our focus from validating input models to validating and making inferences about simulation outputs. Most of this lecture is a review of statistics and reasons for the assumptions for various parametric and non-exact non-parametric methods. We also introduce a few more advanced statistical topics, such as non-parametric methods and special high-power tests for normality. We then switch to focusing on simulations and their outputs, starting with the definition of terminating and non-terminating systems as well as the related transient and steady-state simulations. We will pick up next time with discussing details related to performance measures (and methods) for transient simulations next time and steady-state simulations after that. Our goal was to discuss the difference between point estimation and interval estimation for simulation, but we will hold off to discuss that topic in the next lecture.

View Details

In this lecture, we review statistical fundamentals – such as the origins of the t-test, the meaning of type-I and type-II error (and alternative terminology for both, such as false positive rate and false negative rate) and the connection to statistical power (sensitivity). We review the Receiver Operating Characteristic (ROC) curve and give a qualitative description of where it gets its shape in a hypothesis test. We close with a validation example (from Lecture H) where we use a power analysis on a one-sample t-test to help justify whether we have gathered enough data to trust that a simulation model is a good match for reality when it has a similar mean output performance to the real system.

View Details

During this lecture slot, we start with slides from Lecture G3 (on goodness of fit) that were missed during the previous lecture due to timing. In particular, we review hypothesis testing fundamentals (type-I error, type-II error, statistical power, sensitivity, false positive rate, true negative rate, receiver operating characteristic, ROC, alpha, beta) and then go into examples of using Chi-squared and Kolmogorov–Smirnov tests for goodness of fit for arbitrary distributions. We also introduce Anderson–Darling (for flexibility and higher power) and Shapiro–Wilk (for high-powered normality testing).

We close with where we originally intended to start – with definitions of testing, verification, validation, and calibration. We will pick up from here next time.

View Details

In this lecture, we (nearly) finish our coverage of Input Modeling, where the focus of this lecture is on parameter estimation and assessing goodness of fit. We review input modeling in general and then briefly review fundamentals of hypothesis testing. We discuss type-I error, p-values, type-II error, effect sizes, and statistical power. We discuss the dangers of using p-values at very large sample sizes (where small p-values are not meaningful) and at very small sample sizes (where large p-values are not meaningful). We give some examples of this applied to best-of-7 sports tournaments and voting. We then discuss different shape parameters (including location, scale, and rate), and then introduce summary statistics (sample mean and sample variance) and maximum likelihood estimation (MLE), with an example for a point estimate of the rate of an exponential. We introduce the chi-squared (lower power) and Kolmogorov–Smirnov (KS, high power) tests for goodness of fit, but we will go into them in more detail at the start of the next lecture.

View Details

In this lecture, we continue discussing the choice of input models in stochastic simulation. Here, we pivot from talking about data collection to selection of the broad family of probabilistic distributions that may be a good fit for data. We start with an example where a histogram leads us to introduce additional input models into a flow chart. The rest of the lecture is about choosing models based on physical intuition and the shape of the sampled data (e.g., the shape of histograms). We close with a discussion of probability plots – Q-Q plots and P-P plots, as are used with "fat-pencil tests" – as a good tool for justifying the choice of a family for a certain data set. The next lecture will go over the actual estimation of the parameters for the chosen families and how to quantitatively assess goodness of fit.

View Details

In this lecture, we review pseudo-random number generation and then introduce random-variate generation by way of inverse-transform sampling. In particular, we start with a review of the two most important properties of a pseudo-random number generator (PRNG), uniformity and independence, and discuss statistically rigorous methods for testing for these two properties. For uniformity, we focus on a Chi-square/Chi-squared test for larger numbers of samples and a Kolmogorov–Smirnov (KS) test for smaller numbers of samples. For independence, we discuss autocorrelation tests and runs test, and then we demonstrate a runs above-and-below-the-mean test. We then shift to discussing inverse-transform sampling for continuous random variates and discrete random variates and how the resulting random-variate generators might be implemented in a tool like Rockwell Automation's Arena.

View Details

In this lecture, we introduce the measure-theoretic concept of a random variable (which is neither random nor a variable) and related terms, such as outcomes, events, probability measures, moments, means, etc. Throughout the lecture, we use the metaphor of probability as mass (and thus probability density as mass density, and a mean as a center of mass). This allows us to discuss the "statistical leverage" of outliers in a distribution (i.e., although they happen infrequently, they still have the ability to shift the mean significantly, as in physical leverage). This sets us up to talk about random processes and particular random variables in the next lecture.

View Details

In this lecture, we review basic probability fundamentals (measure spaces, probability measures, random variables, probability density functions, probability mass functions, cumulative distribution functions, moments, mean/expected value/center of mass, standard deviation, variance), and then we start to build a vocabulary of different probabilistic models that are used in different modeling contexts. These include uniform, triangular, normal, exponential, Erlang-k, Weibull, and Poisson variables. If we do not have time to do so during this lecture, we will finish the discussion in the next lecture with the Bernoulli-based discrete variables and Poisson processes.

View Details

In this lecture, we introduce the measure-theoretic concept of a random variable (which is neither random nor a variable) and related terms, such as outcomes, events, probability measures, moments, means, etc. Throughout the lecture, we use the metaphor of probability as mass (and thus probability density as mass density, and a mean as a center of mass). This allows us to discuss the "statistical leverage" of outliers in a distribution (i.e., although they happen infrequently, they still have the ability to shift the mean significantly, as in physical leverage). This sets us up to talk about random processes and particular random variables in the next lecture.

View Details

This lecture (slides embedded below) provides some historical background and motivation for System Dynamics Modeling (SDM) and Agent-Based Modeling (ABM), two other simulation modeling approaches that contrast with Discrete Event System (DES) simulation.

In particular, in this lecture, we briefly introduce System Dynamics Modeling (SDM) and Agent-Based/Individual-Based Modeling (ABM/IBM) as the two ends of the simulation modeling spectrum (from low resolution to high resolution). The introduction of ABM describes applications in life sciences, social sciences, and engineering (Multi-Agent Systems, MAS)/operations research. NetLogo is introduced (as part of preparation for Lab 4), and it is used to present examples of running ABM's as well as the code behind them. This lecture is also coupled with notes discussing the Lab 3 (Monte Carlo simulation) results and general experience. These comments focus on interval estimation (which is right 95% of the time, as opposed to point estimation that is right 0% of the time) and the role of non-trivial distributions of random variables (as opposed to just their means).

View Details

This lecture covers content related to implementing simulations with spreadsheets and the motivations for the use of special-purpose Discrete Event System Simulation tools. In particular, we discuss different approaches to implementing Discrete Event System (DES) simulations (DESS) with simple spreadsheets (e.g., Microsoft Excel, Google Sheets, Apple Numbers, etc.). We cover inventory management problems (such as the newsvendor model) as well as Monte Carlo sampling and stochastic activity networks (SAN's). Although we show that spreadsheets can be very powerful for this kind of work, we highlight that this approach is cumbersome for systems with increasing complexity. So this motivates why we would use more sophisticated tools specifically built for simulation (but perhaps not so great for data analysis by themselves), like Arena, FlexSim, Simio, and NetLogo.

This lecture was recorded by Theodore Pavlic as part of IEE 475 (Simulating Stochastic Systems) at Arizona State University.

View Details

In this lecture, we close out our review of DES fundamentals and hand simulation. After going through a hand-simulation example one last time, we show how to implement a Discrete Event System (DES) simulation using a spreadsheet tool like Microsoft Excel without any "macros" (VBA, etc.). This involves defining relationships ACROSS TIME that allow the spreadsheet to (in a declarative fashion) reconstruct the trajectory that is the output of the simulation.

We then pivot to discussing the previous "Lab 2 (Muffin Oven Simulation)", which lets us introduce common random numbers (CRNs), statistical blocking, requirements of 2-sample and paired t-tests, and more sophisticated statistical methods that better characterize PRACTICAL significance (and take into account the multiple comparisons problem). Thus, the post-lab2 reflections are largely a preview of future topics in the course.

View Details

In this lecture, we review fundamentals of Discrete Event System (DES) simulation (e.g., entities, resources, activities, processes, delays, attributes) and we run through a number of DES modeling examples. These examples show how different research/operations questions can lead to different choices of entities/resources/etc. We close with a hand-simulation example of a single-channel, single-server queue with provided interarrival times and service times.

View Details

In this lecture, we cover fundamentals of discrete-event system (DES) simulation (DESS). This involves reviewing basic simulation concepts (entities, resources, attributes, events, activities, delays) and introducing the event-scheduling world view, which provides a causality framework on which an automatic simulation of a DES system can be built. We also discuss briefly how the stochastic modeling inherent to DESS means that outputs will be variable and thus will require rigorous statistics to make sense of.

View Details

In this lecture, we introduce the three different simulation methodologies (agent-based modeling, system dynamics modeling, and discrete event system simulation) and then focus on how stochastic modeling is used within discrete-event system simulation. In particular, we define terms such as system, dynamic system, state, state variable, activity, delay, resource, entity, and the notion of "input modeling."

View Details

This lecture introduces the topic of modeling with particular focus on the role of quantitative modeling in industrial engineering and operations research. This is an introduction to a course on stochastic simulation.

View Details

In this lecture, we outline the structure and purpose of IEE 475 (Simulating Stochastic Systems) for the Fall 2024 semester at Arizona State University. We go over topics covered in the syllabus and on the course learning management system website.

View Details

In this lecture, we prepare for the final exam and give a brief review of all topics from the course.

View Details

In this lecture, we wrap up the course content in IEE 475. We first do a quick overview of the four variance reduction techniques (VRT's) covered in the previous unit. That is, we cover: common random numbers (CRN's), antithetic variates (AV's), importance sampling, and control variates. We then remember some general comments about the goal of modeling and commonalities seen across simulation platforms (as well as the different types of simulation platforms in general).

View Details

In this lecture, we start by reviewing approaches for absolute and relative performance estimation in stochastic simulation. This begins with a reminder of the use of confidence intervals for estimation of performance for a single simulation model. We then move to different ways to use confidence intervals on mean DIFFERENCES to compare two different simulation models. We then move to the ranking and selection problem for three or more different simulation models, which allows us to talk about analysis of variance (ANOVA) and post hoc tests (like the Tukey HSD or Fisher's LSD). After that review, we move on to introducing variance reduction techniques (VRTs) which reduce the size of confidence intervals by experimentally controlling/accounting for alternative sources of variance (and thus reducing the observed variance in response variables). We discuss Common Random Numbers (CRNs), which use a paired/blocked design to reduce the variance caused by different random-number streams. We start to discuss control variates (CVs), but that discussion will be picked up at the start of the next lecture.

View Details

In this lecture, we review what we have learned about one-sample confidence intervals (i.e., how to use them as graphical versions of one-sample t-tests) for absolute performance estimation in order to motivate the problem of relative performance estimation. We introduce two-sample confidence intervals (i.e., confidence intervals on DIFFERENCES based on different two-sample t-tests) that are tested against a null hypothesis of 0. This means covering confidence interval half widths for the paired-difference t-test, the equal-variance (pooled) t-test, and Welch's unequal variance t-test. Each of these different experimental conditions sets up a different standard error of the mean formula and formula for degrees of freedom that are used to define the actual confidence interval half widths (centered on the difference in sample means in the pairwise comparison of systems). We then generalize to the case of more than 2 systems, particularly for "ranking and selection (R&S)." This lets us review the multiple-comparisons problem (and Bonferroni correction) and how post hoc tests (after an ANOVA) are more statistically powerful ways to do comparisons.

View Details

In this lecture, we start by further reviewing confidence intervals (where they come from and what they mean) and prediction intervals and then use them to motivate a simpler way to determine how many replications are needed in a simulation study (focusing first on transient simulations of terminating systems). We then shift our attention to steady-state simulations of non-terminating systems and the issue of initialization bias. We discuss different methods of "warming up" a steady-state simulation to reduce initialization bias and then merge that discussion with the prior discussion on how to choose the number of replications. In the next lecture, we'll finish up with a discussion of the method of "batch means" in steady-state simulations.

View Details

In this lecture, we review estimating absolute performance from simulation, with focus on choosing the number of necessary replications of transient simulations of terminating systems. The lecture starts by overviewing point estimation, bias, and different types of point estimators. This includes an overview of quantile estimation and how to use quantile estimation to use simulations as null-hypothesis-prediction generators. We the introduce interval estimation with confidence intervals and prediction intervals. Confidence intervals, which are visualizations of t-tests, provide an alternative way to choose the number of required replications without doing a formal power analysis.

View Details

In this lecture, we introduce the estimation of absolute performance measures in simulation – effectively shifting our focus from validating input models to validating and making inferences about simulation outputs. Most of this lecture is a review of statistics and reasons for the assumptions for various parametric and non-exact non-parametric methods. We also introduce a few more advanced statistical topics, such as non-parametric methods and special high-power tests for normality. We then switch to focusing on simulations and their outputs, starting with the definition of terminating and non-terminating systems as well as the related transient and steady-state simulations. We will pick up next time with discussing details related to performance measures (and methods) for transient simulations next time and steady-state simulations after that. Our goal was to discuss the difference between point estimation and interval estimation for simulation, but we will hold off to discuss that topic in the next lecture.

View Details

In this lecture, we review statistical fundamentals – such as the origins of the t-test, the meaning of type-I and type-II error (and alternative terminology for both, such as false positive rate and false negative rate) and the connection to statistical power (sensitivity). We review the Receiver Operating Characteristic (ROC) curve and give a qualitative description of where it gets its shape in a hypothesis test. We close with a validation example (from the previous lecture) where we use a power analysis on a one-sample t-test to help justify whether we have gathered enough data to trust that a simulation model is a good match for reality when it has a similar mean output performance to the real system.

View Details

In this lecture, we mostly cover slides from Lecture G3 (on goodness of fit) that were missed during the previous lecture. In particular, we review hypothesis testing fundamentals (type-I error, type-II error, statistical power, sensitivity, false positive rate, true negative rate, receiver operating characteristic, ROC, alpha, beta) and then go into examples of using Chi-squared and Kolmogorov–Smirnov tests for goodness of fit for arbitrary distributions. We also introduce Anderson–Darling (for flexibility and higher power) and Shapiro–Wilk (for high-powered normality testing). We close with where we originally intended to start – with definitions of testing, verification, validation, and calibration. We will pick up from here next time.

View Details

In this lecture, we (nearly) finish our coverage of Input Modeling, where the focus of this lecture is on parameter estimation and assessing goodness of fit. We review input modeling in general and then briefly review fundamentals of hypothesis testing. We discuss type-I error, p-values, type-II error, effect sizes, and statistical power. We discuss the dangers of using p-values at very large sample sizes (where small p-values are not meaningful) and at very small sample sizes (where large p-values are not meaningful). We give some examples of this applied to best-of-7 sports tournaments and voting. We then discuss different shape parameters (including location, scale, and rate), and then introduce summary statistics (sample mean and sample variance) and maximum likelihood estimation (MLE), with an example for a point estimate of the rate of an exponential. We introduce the chi-squared (lower power) and Kolmogorov–Smirnov (KS, high power) tests for goodness of fit, but we will go into them in more detail at the start of the next lecture.

View Details

In this lecture, we continue discussing the choice of input models in stochastic simulation. Here, we pivot from talking about data collection to selection of the broad family of probabilistic distributions that may be a good fit for data. We start with an example where a histogram leads us to introduce additional input models into a flow chart. The rest of the lecture is about choosing models based on physical intuition and the shape of the sampled data (e.g., the shape of histograms). We close with a discussion of probability plots – Q-Q plots and P-P plots, as are used with "fat-pencil tests" – as a good tool for justifying the choice of a family for a certain data set. The next lecture will go over the actual estimation of the parameters for the chosen families and how to quantitatively assess goodness of fit.

View Details

In this lecture, we introduce the detailed process of input modeling. Input models are probabilistic models that introduce variation in simulation models of systems. Those input models must be chosen to match statistical distributions in data. Over this unit, we cover collection of data for this process, choice of probabilistic families to fit to these data, and then optimized parameter choice within those families and evaluation of fit with goodness of fit. In this lecture, we discuss issues related to data collection.

View Details

Midterm review session for ASU IEE 475 for Fall 2022.

Whiteboard notes for this lecture can be found at: https://www.dropbox.com/s/ljc61rarhns41u2/2022-Fall-Midterm_Review_Notes.pdf?dl=0

View Details

In this lecture, we review pseudo-random number generation and then introduce random-variate generation by way of inverse-transform sampling. In particular, we start with a review of the two most important properties of a pseudo-random number generator (PRNG), uniformity and independence, and discuss statistically rigorous methods for testing for these two properties. For uniformity, we focus on a Chi-square/Chi-squared test for larger numbers of samples and a Kolmogorov–Smirnov (KS) test for smaller numbers of samples. For independence, we discuss autocorrelation tests and runs test, and then we demonstrate a runs above-and-below-the-mean test. We then shift to discussing inverse-transform sampling for continuous random variates and discrete random variates and how the resulting random-variate generators might be implemented in a tool like Rockwell Automation's Arena.

View Details

In this lecture, we first cover some discrete distributions (and the Poisson process) that we ran out of time for during the previous lecture. We then launch into a discussion of how to generate pseudo-random numbers distributed uniformly between 0 and 1 (which are necessary for us to easily generate random variates of any distribution). We talk about the two most important properties of a pseudo-random number generator (PRNG), uniformity and independence. We then talk about desirable properties. Some examples are given of some early PRNG's, and then we introduce the linear congruential generator (LCG) and its variants (like the Combined Linear Congruential Generator, CLCG), which represent a much more modern PRNG that has a number of good properties. We close with a discussion of tests of uniformity. We will continue this discussion and add on tests for independence during next lecture (which will primarily cover random-VARIATE generation).

View Details

In this lecture, we review basic probability fundamentals (measure spaces, probability measures, random variables, probability density functions, probability mass functions, cumulative distribution functions, moments, mean/expected value/center of mass, standard deviation, variance), and then we start to build a vocabulary of different probabilistic models that are used in different modeling contexts. These include uniform, triangular, normal, exponential, Erlang-k, Weibull, and Poisson variables. We will finish the discussing next time with the Bernoulli-based discrete variables and Poisson processes.

View Details

In this lecture, we introduce the measure-theoretic concept of a random variable (which is neither random nor a variable) and related terms, such as outcomes, events, probability measures, moments, means, etc. Throughout the lecture, we use the metaphor of probability as mass (and thus probability density as mass density, and a mean as a center of mass). This allows us to discuss the "statistical leverage" of outliers in a distribution (i.e., although they happen infrequently, they still have the ability to shift the mean significantly, as in physical leverage). This sets us up to talk about random processes and particular random variables in the next lecture.

View Details

In this lecture, we briefly introduce System Dynamics Modeling (SDM) and Agent-Based/Individual-Based Modeling (ABM/IBM) as the two ends of the simulation modeling spectrum (from low resolution to high resolution). The introduction of ABM describes applications in life sciences, social sciences, and engineering (Multi-Agent Systems, MAS)/operations research. NetLogo is introduced, and it is used to present examples of running ABM's as well as the code behind them. At the end of the ABM/NetLogo introduction, comments about the previous lab on Monte Carlo simulation are given. These comments focus on interval estimation (which is right 95% of the time, as opposed to point estimation that is right 0% of the time) and the role of non-trivial distributions of random variables (as opposed to just their means).

View Details

In this lecture, we discuss different approaches to implementing Discrete Event System (DES) simulations (DESS) with simple spreadsheets (e.g., Microsoft Excel, Google Sheets, Apple Numbers, etc.). We cover inventory management problems (such as the newsvendor model) as well as Monte Carlo sampling and stochastic activity networks (SAN's). Although we show that spreadsheets can be very powerful for this kind of work, we highlight that this approach is cumbersome for systems with increasing complexity. So this motivates why we would use more sophisticated tools specifically built for simulation (but perhaps not so great for data analysis by themselves), like Arena, FlexSim, Simio, and NetLogo.

View Details

In this lecture, we close out our review of DES fundamentals and hand simulation. After going through a hand-simulation example one last time, we show how to implement a Discrete Event System (DES) simulation using a spreadsheet tool like Microsoft Excel without any "macros" (VBA, etc.). This involves defining relationships ACROSS TIME that allow the spreadsheet to (in a declarative fashion) reconstruct the trajectory that is the output of the simulation. We then pivot to discussing the previous "Lab 2 (Muffin Oven Simulation)", which lets us discuss common random numbers (CRNs), statistical blocking, requirements of 2-sample and paired t-tests, and more sophisticated statistical methods that better characterize PRACTICAL significance (and take into account the multiple comparisons problem). Thus, the post-lab2 reflections are largely a preview of future topics in the course.

View Details

In this lecture, we review fundamentals of Discrete Event System (DES) simulation (e.g., entities, resources, activities, processes, delays, attributes) and we run through a number of DES modeling examples. These examples show how different research/operations questions can lead to different choices of entities/resources/etc. We close with a hand-simulation example of a single-channel, single-server queue with provided interarrival times and service times.

View Details

In this lecture, we cover fundamentals of discrete-event system (DES) simulation (DESS). This involves reviewing basic simulation concepts (entities, resources, attributes, events, activities, delays) and introducing the event-scheduling world view, which provides a causality framework on which an automatic simulation of a DES system can be built. We also discuss briefly how the stochastic modeling inherent to DESS means that outputs will be variable and thus will require rigorous statistics to make sense of.

View Details

In this lecture, we introduce the three different simulation methodologies (agent-based modeling, system dynamics modeling, and discrete event system simulation) and then focus on how stochastic modeling is used within discrete-event system simulation.

View Details

In this lecture, we introduce Industrial and Systems Engineering as a blend of science and engineering that necessitates model building. We then define model (as something that answers a "What If" question) and different types of models. This gives us an opportunity to discuss how modeling is less about describing reality and more about generating tools to do useful things/make useful predictions. We end with a comparison of mental and quantitative models, as well as a comparison of different types of quantitative models (including simulation modeling).

View Details

In this lecture, we go over course policies for the Fall 2022 session of IEE 475.

View Details

This lecture section is a cumulative review of material from the semester and is meant to serve as a study guide for students preparing for the upcoming final exam. Topics start at modeling fundamentals (what is the purpose of a model in general) to the specifics of designing statistical experiments with stochastic simulations.

[ due to an instructor error, the lecture from 2021-11-09 was not recorded, and the archived 2020-12-01 lecture is re-used here instead ]

View Details

In this lecture, we wrap up our discussion of Variance Reduction Techniques. We introduced Common Random Numbers (CRNs) last time, which we review in this lecture. We then introduce Control Variates (CVs), Antithetic Variates (AVs), and Importance Sampling. These four methods are all examples of amplifying signals in a statistical experiment either by manipulating the simulation execution or using information about known sources of variance to increase statistical power.

View Details

In this lecture, we wrap up our discussion of the movement from point estimation (sample means) to interval estimation for: (a) estimating absolute performance of a system, (b) estimating relative performance of two systems, and (c) estimating relative performance of more than 2 systems. We then pivot to discussing Variance Reduction Techniques (VRT's), starting with Common Random Numbers (CRN's).

View Details

In this lecture, we further review the use of confidence intervals to summarize empirical results from simulation as we move from thinking about absolute performance estimation (i.e., using one model system to estimate one parameter) to relative performance estimation (i.e., comparing two model systems to make an inference about whether they differ). This allows us to discuss how confidence intervals are used in regression analysis and start to motivate how to build confidence intervals that are summarizes of 2-sample (instead of 1-sample) t-tests. We had to stop a little early, and so the next lecture will discuss how to convert paired t-tests and two different types of 2-sample t-tests into 2-sample confidence intervals (which are compared to 0).

View Details

This lecture continues to discuss issues related to estimating absolute performance from transient and steady-state simulations (of terminating and non-terminating systems, respectively). We continue to emphasize the importance and utility of interval estimations (over point estimates). We then move on to discuss experimental methodologies useful for steady-state simulations, particularly related to eliminating estimator bias and reducing computational time.

[ due to an instructor error, the lecture from 2021-11-09 was not recorded, and the archived 2020-11-010lecture is re-used here instead ]

View Details

In this lecture, we continue to introduce terminating and non-terminating systems and difference methods for estimating performance from simulation models of them (using transient and steady-state simulations). This involves a description of various types of point estimators (mean and quantile) as well as related interval estimators (confidence intervals and prediction intervals, as well as the relationship to standard error of the mean (SEM)). We start to discuss issues involving making inferences from pseudo-replicated within-replication samples versus across-replication samples (which are independent and often normally distributed). We will continue this in the next lecture, as we start focusing more on steady-state simulations of non-terminating systems.

[ For some strange reason, the in-room video camera was not recorded as the speaker view despite apparently working during the class. Consequently, only the slide view is shown. ]

View Details

In this lecture, we review the fundamental tradeoffs in hypothesis testing and the concrete origins of the assumptions in both the t-test and Chi-square test. We also discuss parametric and non-parametric statistics (including exact and non-exact tests) and how non-parametric, exact statistics like the Kolmogorov–Smirnov test are derived. This culminates in a discussion of the multiple comparisons (MC) problem and the Bonferroni correction as well as alternative tests (such as a MANOVA or an ANOVA with post hoc test) that have more statistical power than the Bonferroni correction. We close with an introduction to performance inference from simulation, which we will continue discussing in the next 3 lectures.

View Details

In this halloween-themed lecture, we go into more detail on the foundations of hypothesis testing – specifically hypothesis testing with small sample sizes. This allows us to talk about where the Student's t test comes from (and why it is defined that way) as well as where the Chi-square test comes from (and why it is defined that way). Throughout the lecture, we highlight the importance of statistical power and do a power analysis example for a paired-difference t-test.

View Details

In this lecture, we review summary statistics, MLE, and goodness-of-fit tests (particularly Chi-square and Kolmogorov–Smirnov, with some mention of Anderson–Darling and Shapiro–Wilk), with a particular focus on the type-I error, type-II error, and statistical power. We then introduce verification, validation, and calibration of simulation models and close with an example for the simulation of a bank. We use rigorous statistical methods to drive the calibration process that leads to updating the model of the bank and ensuring its outputs are a good statistical match for outputs in a real bank. This involves making use of a power analysis for a one-sample, two-sided t-test. We will cover the paired t-test version of this problem in the next lecture.

View Details

In this lecture, we start out with Q-Q and P-P probability plots that we did not have time to cover from last time. We then transition to a review about type-I error and p values and try to motivate the topics of STATISTICAL POWER and EFFECT SIZES, which we will dive into more in the next few lectures. We then discuss summary statistics and how to use methods such as maximum likelihood estimation (MLE) to come up with good choices of parameters for distributions picked in the input modeling process. Next time, we will discuss testing the (goodness of) fit for those parameterized distributions.

View Details

In this lecture, we continue our discussion of input modeling in depth. We start with a more detailed example of how data collection can guide the choice of the structural features of a system. We then move to the point in the process when the structure of the model is set but the input models have to be chosen based on collected data. We cover methods for generating histograms and matching those histograms to common distributions (both discrete and continuous). We stop just before discussing Q-Q plots and P-P plots, which we will pick up next time along with discussing how to parameterize these chosen distributions.

View Details

In this lecture, we introduce the 3-lecture unit on "Input Modeling." We start with motivations from thinking about stochastic simulation models and then describe the potential problems that can occur in collecting data. We close with a set of rules that can be helpful to follow when collecting data. We will start on choosing probabilistic families, parameterizing them, and testing goodness of fit next lecture (and extending over the next lecture).

View Details

In this lecture, we review topics from the first half of the semester that will be tested over in the upcoming midterm. Most of the class involves working examples on the whiteboard.

Whiteboard notes captured for this session can be found at: https://www.dropbox.com/s/pih0wt3abwbatbb/IEE475-LectureF-2021-09-30-Midterm_Review-Whiteboard_Notes.pdf?dl=0

View Details

In this lecture, we finish covering tests of uniformity (Chi-squared and Kolmogorov–Smirnov) and independence (autocorrelation and runs (above and below) tests) for pseudo-random number generators (PRNGs). We then move on to discussing the details of inverse-transform sampling for random-variate generation. We cover how to derive a CDF from a piecewise PDF and how to invert a CDF to produce a quantile function fit for random-variate generation. We also discuss the discrete inverse-transform case.

View Details

We start the lecture covering some discrete random variables that we did not get to during Lecture D2. We also introduce the Poisson process and how it relates to the Poisson and exponential random variables. We then pivot to discussing pseudo-random number generators (PRNGs), including their required as well as desired properties and statistical tests to test for independence and uniformity. We will continue the discussion of statistical tests for independence at the start of next lecture (Lecture E2).

View Details

In this lecture, we review basic probability space concepts from the previous lecture. We then go on to discuss the common probabilistic models that we will use in stochastic simulation (e.g., uniform, triangular, normal, exponential, Weibull, Erlang, Poisson, etc.). Basic background on the structure of each distribution is given as well as practical reasons why one distribution might be picked over another.

View Details

In this lecture, we use motivation from stochastic modeling (i.e., incorporating randomness into models in order to capture realistic variation without having to specify a great many details) to formally introduce random variables and probability spaces (as a subset of measure theory). We heavily lean on the analogy between probability and mass as we introduce the sample space, probability measure, random variable, probability mass function (pmf), probability density function (pdf), cumulative distribution function (cdf), and moments (including expectation and central moments as in variance).

View Details

In this lecture, we review results from the Monte Carlo simulation lab (Lab 3) and setup motivation for the agent-based modeling/NetLogo lab (Lab 4). For the MC-lab review, we cover the estimation of pi by drawing random coordinates in the unit cube. We also discuss the possibly counter-intuitive results from estimating the length of a 3-path stochastic activity network. To prepare for Lab 4, we review the three different types of simulation methodologies (ABM/IBM, DES, and SDM) and then give a brief introduction to NetLogo. A more detailed/tutorial introduction to NetLogo will take place during Lab 4.

View Details

In this lecture, we discuss more sophisticated dynamical simulation models that can be implemented within spreadsheets. We start with a review of the M/M/1 single-channel, single-server queueing node and then show how more explicit state variables can be introduced in an M/M/2 version (i.e., with two servers). We then discuss two different popular inventory management models (implemented within a spreadsheet) -- the "Order-up-to (M,N)" model as well as the "newsvendor (single-period/perishable) model". We close with some discussion of Monte Carlo methods -- which apply simulation techniques as numerical methods to solve mathematical problems that might otherwise be intractable analytically. Despite all of these examples of the power of spreadsheets, we end with a hint that much more is possible in terms of simulation of complex systems if we use specialized simulation tools. We will introduce some of those more specialized tools starting in the next lecture.

View Details

In this lecture, we review hand-simulation/DES simulation basics. We then introduce how to simulate discrete event system simulations (which are dynamic simulation models built around the idea of "state") in declarative programming frameworks like spreadsheets (which have no "state"). We work through the relationships necessary to encode in a spreadsheet to simulate a single-channel, single-server queue. We then pivot to covering comments from Lab 2, which was a hand simulation of a system with partial batching. This allows for motivating why we use multiple replications when we do empirical work with stochastic simulations, and how tools such as common random numbers can reduce variance but require special statistical tools (such as the paired-difference t-test). We then discuss the multiple comparisons problem, some ways to solve it, and how linear models extend what we can say from empirical studies -- so we can go from statistically significant to practically significant. (i.e., we can better characterize the "effect size" of a variable we have control over).

View Details

In this lecture, we carry forward our high-level description of the event-scheduling world view to specific hand-simulation examples of a single-channel, single-server queueing network node.

View Details

In this lecture, we review modeling basics for process-centric modeling (entities, resources, events, activities, delays, etc.) and then introduce the event-scheduling world view that acts behind the scenes in any discrete event system (DES) simulation. We begin discussing hand simulation of DESS, at least in the abstract. More concrete examples are to come in the next lecture.

View Details

In this lecture, we pivot from our general introduction to (quantitative) modeling to a more specific introduction of simulation modeling. System dynamics modeling (SDM), agent-based modeling (ABM), and discrete event system (DES) simulation are introduced, with the most detail on DES that will be the focus for the course. We then motivate the approach of "stochastic modeling" -- using randomness in these models in place of deterministic details.

View Details

In this lecture, we introduce the basic motivations for quantitative modeling -- including fundamental definitions of what is a model. This definition is meant to cover all models -- from fashion models to mouse models to statistical models to simulation models.

View Details

Recorded day-1 lecture of IEE 475 (Simulating Stochastic Systems) in the Fall 2021 semester. Introduces course and its policies. Audio is poor due to microphone support in room.

Pre-recorded versions of both parts of the lecture above with much better audio (and video):