Stochastic process
Stochastic process
Main page
2310268

Stochastic process

logo
Community Hub0 subscribers
Read side by side
from Wikipedia
A computer-simulated realization of a Wiener or Brownian motion process on the surface of a sphere. The Wiener process is widely considered the most studied and central stochastic process in probability theory.[1][2][3]

In probability theory and related fields, a stochastic (/stəˈkæstɪk/) or random process is a mathematical object usually defined as a family of random variables in a probability space, where the index of the family often has the interpretation of time. Stochastic processes are widely used as mathematical models of systems and phenomena that appear to vary in a random manner. Examples include the growth of a bacterial population, an electrical current fluctuating due to thermal noise, or the movement of a gas molecule.[1][4][5] Stochastic processes have applications in many disciplines such as biology,[6] chemistry,[7] ecology,[8] neuroscience,[9] physics,[10] image processing, signal processing,[11] control theory,[12] information theory,[13] computer science,[14] and telecommunications.[15] Furthermore, seemingly random changes in financial markets have motivated the extensive use of stochastic processes in finance.[16][17][18]

Applications and the study of phenomena have in turn inspired the proposal of new stochastic processes. Examples of such stochastic processes include the Wiener process or Brownian motion process,[a] used by Louis Bachelier to study price changes on the Paris Bourse,[21] and the Poisson process, used by A. K. Erlang to study the number of phone calls occurring in a certain period of time.[22] These two stochastic processes are considered the most important and central in the theory of stochastic processes,[1][4][23] and were invented repeatedly and independently, both before and after Bachelier and Erlang, in different settings and countries.[21][24]

The term random function is also used to refer to a stochastic or random process,[25][26] because a stochastic process can also be interpreted as a random element in a function space.[27][28] The terms stochastic process and random process are used interchangeably, often with no specific mathematical space for the set that indexes the random variables.[27][29] But often these two terms are used when the random variables are indexed by the integers or an interval of the real line.[5][29] If the random variables are indexed by the Cartesian plane or some higher-dimensional Euclidean space, then the collection of random variables is usually called a random field instead.[5][30] The values of a stochastic process are not always numbers and can be vectors or other mathematical objects.[5][28]

Based on their mathematical properties, stochastic processes can be grouped into various categories, which include random walks,[31] martingales,[32] Markov processes,[33] Lévy processes,[34] Gaussian processes,[35] random fields,[36] renewal processes, and branching processes.[37] The study of stochastic processes uses mathematical knowledge and techniques from probability, calculus, linear algebra, set theory, and topology[38][39][40] as well as branches of mathematical analysis such as real analysis, measure theory, Fourier analysis, and functional analysis.[41][42][43] The theory of stochastic processes is considered to be an important contribution to mathematics[44] and it continues to be an active topic of research for both theoretical reasons and applications.[45][46][47]

Introduction

[edit]

A stochastic or random process can be defined as a collection of random variables that is indexed by some mathematical set, meaning that each random variable of the stochastic process is uniquely associated with an element in the set.[4][5] The set used to index the random variables is called the index set. Historically, the index set was some subset of the real line, such as the natural numbers, giving the index set the interpretation of time.[1] Each random variable in the collection takes values from the same mathematical space known as the state space. This state space can be, for example, the integers, the real line or -dimensional Euclidean space.[1][5] An increment is the amount that a stochastic process changes between two index values, often interpreted as two points in time.[48][49] A stochastic process can have many outcomes, due to its randomness, and a single outcome of a stochastic process is called, among other names, a sample function or realization.[28][50]

A single computer-simulated sample function or realization, among other terms, of a three-dimensional Wiener or Brownian motion process for time 0 ≤ t ≤ 2. The index set of this stochastic process is the non-negative numbers, while its state space is three-dimensional Euclidean space.

Classifications

[edit]

A stochastic process can be classified in different ways, for example, by its state space, its index set, or the dependence among the random variables. One common way of classification is by the cardinality of the index set and the state space.[51][52][53]

When interpreted as time, if the index set of a stochastic process has a finite or countable number of elements, such as a finite set of numbers, the set of integers, or the natural numbers, then the stochastic process is said to be in discrete time.[54][55] If the index set is some interval of the real line, then time is said to be continuous. The two types of stochastic processes are respectively referred to as discrete-time and continuous-time stochastic processes.[48][56][57] Discrete-time stochastic processes are considered easier to study because continuous-time processes require more advanced mathematical techniques and knowledge, particularly due to the index set being uncountable.[58][59] If the index set is the integers, or some subset of them, then the stochastic process can also be called a random sequence.[55]

If the state space is the integers or natural numbers, then the stochastic process is called a discrete or integer-valued stochastic process. If the state space is the real line, then the stochastic process is referred to as a real-valued stochastic process or a process with continuous state space. If the state space is -dimensional Euclidean space, then the stochastic process is called a -dimensional vector process or -vector process.[51][52]

Etymology

[edit]

The word stochastic in English was originally used as an adjective with the definition "pertaining to conjecturing", and stemming from a Greek word meaning "to aim at a mark, guess", and the Oxford English Dictionary gives the year 1662 as its earliest occurrence.[60] In his work on probability Ars Conjectandi, originally published in Latin in 1713, Jakob Bernoulli used the phrase "Ars Conjectandi sive Stochastice", which has been translated to "the art of conjecturing or stochastics".[61] This phrase was used, with reference to Bernoulli, by Ladislaus Bortkiewicz[62] who in 1917 wrote in German the word stochastik with a sense meaning random. The term stochastic process first appeared in English in a 1934 paper by Joseph Doob.[60] For the term and a specific mathematical definition, Doob cited another 1934 paper, where the term stochastischer Prozeß was used in German by Aleksandr Khinchin,[63][64] though the German term had been used earlier, for example, by Andrei Kolmogorov in 1931.[65]

According to the Oxford English Dictionary, early occurrences of the word random in English with its current meaning, which relates to chance or luck, date back to the 16th century, while earlier recorded usages started in the 14th century as a noun meaning "impetuosity, great speed, force, or violence (in riding, running, striking, etc.)". The word itself comes from a Middle French word meaning "speed, haste", and it is probably derived from a French verb meaning "to run" or "to gallop". The first written appearance of the term random process pre-dates stochastic process, which the Oxford English Dictionary also gives as a synonym, and was used in an article by Francis Edgeworth published in 1888.[66]

Terminology

[edit]

The definition of a stochastic process varies,[67] but a stochastic process is traditionally defined as a collection of random variables indexed by some set.[68][69] The terms random process and stochastic process are considered synonyms and are used interchangeably, without the index set being precisely specified.[27][29][30][70][71][72] Both "collection",[28][70] or "family" are used[4][73] while instead of "index set", sometimes the terms "parameter set"[28] or "parameter space"[30] are used.

The term random function is also used to refer to a stochastic or random process,[5][74][75] though sometimes it is only used when the stochastic process takes real values.[28][73] This term is also used when the index sets are mathematical spaces other than the real line,[5][76] while the terms stochastic process and random process are usually used when the index set is interpreted as time,[5][76][77] and other terms are used such as random field when the index set is -dimensional Euclidean space or a manifold.[5][28][30]

Notation

[edit]

A stochastic process can be denoted, among other ways, by ,[56] ,[69] [78] or simply as . Some authors mistakenly write even though it is an abuse of function notation.[79] For example, or are used to refer to the random variable with the index , and not the entire stochastic process.[78] If the index set is , then one can write, for example, to denote the stochastic process.[29]

Examples

[edit]

Bernoulli process

[edit]

One of the simplest stochastic processes is the Bernoulli process,[80] which is a sequence of independent and identically distributed (iid) random variables, where each random variable takes either the value one or zero, say one with probability and zero with probability . This process can be linked to an idealisation of repeatedly flipping a coin, where the probability of obtaining a head is taken to be and its value is one, while the value of a tail is zero.[81] In other words, a Bernoulli process is a sequence of iid Bernoulli random variables,[82] where each idealised coin flip is an example of a Bernoulli trial.[83]

Random walk

[edit]

Random walks are stochastic processes that are usually defined as sums of iid random variables or random vectors in Euclidean space, so they are processes that change in discrete time.[84][85][86][87][88] But some also use the term to refer to processes that change in continuous time,[89] particularly the Wiener process used in financial models, which has led to some confusion, resulting in its criticism.[90] There are various other types of random walks, defined so their state spaces can be other mathematical objects, such as lattices and groups, and in general they are highly studied and have many applications in different disciplines.[89][91]

A classic example of a random walk is known as the simple random walk, which is a stochastic process in discrete time with the integers as the state space, and is based on a Bernoulli process, where each Bernoulli variable takes either the value positive one or negative one. In other words, the simple random walk takes place on the integers, and its value increases by one with probability, say, , or decreases by one with probability , so the index set of this random walk is the natural numbers, while its state space is the integers. If , this random walk is called a symmetric random walk.[92][93]

Wiener process

[edit]

The Wiener process is a stochastic process with stationary and independent increments that are normally distributed based on the size of the increments.[2][94] The Wiener process is named after Norbert Wiener, who proved its mathematical existence, but the process is also called the Brownian motion process or just Brownian motion due to its historical connection as a model for Brownian movement in liquids.[95][96][97]

Realizations of Wiener processes (or Brownian motion processes) with drift (blue) and without drift (red)

Playing a central role in the theory of probability, the Wiener process is often considered the most important and studied stochastic process, with connections to other stochastic processes.[1][2][3][98][99][100][101] Its index set and state space are the non-negative numbers and real numbers, respectively, so it has both continuous index set and states space.[102] But the process can be defined more generally so its state space can be -dimensional Euclidean space.[91][99][103] If the mean of any increment is zero, then the resulting Wiener or Brownian motion process is said to have zero drift. If the mean of the increment for any two points in time is equal to the time difference multiplied by some constant , which is a real number, then the resulting stochastic process is said to have drift .[104][105][106]

Almost surely, a sample path of a Wiener process is continuous everywhere but nowhere differentiable. It can be considered as a continuous version of the simple random walk.[49][105] The process arises as the mathematical limit of other stochastic processes such as certain random walks rescaled,[107][108] which is the subject of Donsker's theorem or invariance principle, also known as the functional central limit theorem.[109][110][111]

The Wiener process is a member of some important families of stochastic processes, including Markov processes, Lévy processes and Gaussian processes.[2][49] The process also has many applications and is the main stochastic process used in stochastic calculus.[112][113] It plays a central role in quantitative finance,[114][115] where it is used, for example, in the Black–Scholes–Merton model.[116] The process is also used in different fields, including the majority of natural sciences as well as some branches of social sciences, as a mathematical model for various random phenomena.[3][117][118]

Poisson process

[edit]

The Poisson process is a stochastic process that has different forms and definitions.[119][120] It can be defined as a counting process, which is a stochastic process that represents the random number of points or events up to some time. The number of points of the process that are located in the interval from zero to some given time is a Poisson random variable that depends on that time and some parameter. This process has the natural numbers as its state space and the non-negative numbers as its index set. This process is also called the Poisson counting process, since it can be interpreted as an example of a counting process.[119]

If a Poisson process is defined with a single positive constant, then the process is called a homogeneous Poisson process.[119][121] The homogeneous Poisson process is a member of important classes of stochastic processes such as Markov processes and Lévy processes.[49]

The homogeneous Poisson process can be defined and generalized in different ways. It can be defined such that its index set is the real line, and this stochastic process is also called the stationary Poisson process.[122][123] If the parameter constant of the Poisson process is replaced with some non-negative integrable function of , the resulting process is called an inhomogeneous or nonhomogeneous Poisson process, where the average density of points of the process is no longer constant.[124] Serving as a fundamental process in queueing theory, the Poisson process is an important process for mathematical models, where it finds applications for models of events randomly occurring in certain time windows.[125][126]

Defined on the real line, the Poisson process can be interpreted as a stochastic process,[49][127] among other random objects.[128][129] But then it can be defined on the -dimensional Euclidean space or other mathematical spaces,[130] where it is often interpreted as a random set or a random counting measure, instead of a stochastic process.[128][129] In this setting, the Poisson process, also called the Poisson point process, is one of the most important objects in probability theory, both for applications and theoretical reasons.[22][131] But it has been remarked that the Poisson process does not receive as much attention as it should, partly due to it often being considered just on the real line, and not on other mathematical spaces.[131][132]

Definitions

[edit]

Stochastic process

[edit]

A stochastic process is defined as a collection of random variables defined on a common probability space , where is a sample space, is a -algebra, and is a probability measure; and the random variables, indexed by some set , all take values in the same mathematical space , which must be measurable with respect to some -algebra .[28]

In other words, for a given probability space and a measurable space , a stochastic process is a collection of -valued random variables, which can be written as:[80]

Historically, in many problems from the natural sciences a point had the meaning of time, so is a random variable representing a value observed at time .[133] A stochastic process can also be written as to reflect that it is actually a function of two variables, and .[28][134]

There are other ways to consider a stochastic process, with the above definition being considered the traditional one.[68][69] For example, a stochastic process can be interpreted or defined as a -valued random variable, where is the space of all the possible functions from the set into the space .[27][68] However this alternative definition as a "function-valued random variable" in general requires additional regularity assumptions to be well-defined.[135]

Index set

[edit]

The set is called the index set[4][51] or parameter set[28][136] of the stochastic process. Often this set is some subset of the real line, such as the natural numbers or an interval, giving the set the interpretation of time.[1] In addition to these sets, the index set can be another set with a total order or a more general set,[1][54] such as the Cartesian plane or -dimensional Euclidean space, where an element can represent a point in space.[48][137] That said, many results and theorems are only possible for stochastic processes with a totally ordered index set.[138]

State space

[edit]

The mathematical space of a stochastic process is called its state space. This mathematical space can be defined using integers, real lines, -dimensional Euclidean spaces, complex planes, or more abstract mathematical spaces. The state space is defined using elements that reflect the different values that the stochastic process can take.[1][5][28][51][56]

Sample function

[edit]

A sample function is a single outcome of a stochastic process, so it is formed by taking a single possible value of each random variable of the stochastic process.[28][139] More precisely, if is a stochastic process, then for any point , the mapping

is called a sample function, a realization, or, particularly when is interpreted as time, a sample path of the stochastic process .[50] This means that for a fixed , there exists a sample function that maps the index set to the state space .[28] Other names for a sample function of a stochastic process include trajectory, path function[140] or path.[141]

Increment

[edit]

An increment of a stochastic process is the difference between two random variables of the same stochastic process. For a stochastic process with an index set that can be interpreted as time, an increment is how much the stochastic process changes over a certain time period. For example, if is a stochastic process with state space and index set , then for any two non-negative numbers and such that , the difference is a -valued random variable known as an increment.[48][49] When interested in the increments, often the state space is the real line or the natural numbers, but it can be -dimensional Euclidean space or more abstract spaces such as Banach spaces.[49]

Further definitions

[edit]

Law

[edit]

For a stochastic process defined on the probability space , the law of stochastic process is defined as the pushforward measure:

where is a probability measure, the symbol denotes function composition and is the pre-image of the measurable function or, equivalently, the -valued random variable , where is the space of all the possible -valued functions of , so the law of a stochastic process is a probability measure.[27][68][142][143]

For a measurable subset of , the pre-image of gives

so the law of a can be written as:[28]

The law of a stochastic process or a random variable is also called the probability law, probability distribution, or the distribution.[133][142][144][145][146]

Finite-dimensional probability distributions

[edit]

For a stochastic process with law , its finite-dimensional distribution for is defined as:

This measure is the joint distribution of the random vector ; it can be viewed as a "projection" of the law onto a finite subset of .[27][147]

For any measurable subset of the -fold Cartesian power , the finite-dimensional distributions of a stochastic process can be written as:[28]

The finite-dimensional distributions of a stochastic process satisfy two mathematical conditions known as consistency conditions.[57]

Stationarity

[edit]

Stationarity is a mathematical property that a stochastic process has when all the random variables of that stochastic process are identically distributed. In other words, if is a stationary stochastic process, then for any the random variable has the same distribution, which means that for any set of index set values , the corresponding random variables

all have the same probability distribution. The index set of a stationary stochastic process is usually interpreted as time, so it can be the integers or the real line.[148][149] But the concept of stationarity also exists for point processes and random fields, where the index set is not interpreted as time.[148][150][151]

When the index set can be interpreted as time, a stochastic process is said to be stationary if its finite-dimensional distributions are invariant under translations of time. This type of stochastic process can be used to describe a physical system that is in steady state, but still experiences random fluctuations.[148] The intuition behind stationarity is that as time passes the distribution of the stationary stochastic process remains the same.[152] A sequence of random variables forms a stationary stochastic process only if the random variables are identically distributed.[148]

A stochastic process with the above definition of stationarity is sometimes said to be strictly stationary, but there are other forms of stationarity. One example is when a discrete-time or continuous-time stochastic process is said to be stationary in the wide sense, then the process has a finite second moment for all and the covariance of the two random variables and depends only on the number for all .[152][153] Khinchin introduced the related concept of stationarity in the wide sense, which has other names including covariance stationarity or stationarity in the broad sense.[153][154]

Filtration

[edit]

A filtration is an increasing sequence of sigma-algebras defined in relation to some probability space and an index set that has some total order relation, such as in the case of the index set being some subset of the real numbers. More formally, if a stochastic process has an index set with a total order, then a filtration , on a probability space is a family of sigma-algebras such that for all , where and denotes the total order of the index set .[51] With the concept of a filtration, it is possible to study the amount of information contained in a stochastic process at , which can be interpreted as time .[51][155] The intuition behind a filtration is that as time passes, more and more information on is known or available, which is captured in , resulting in finer and finer partitions of .[156][157]

Modification

[edit]

A modification of a stochastic process is another stochastic process, which is closely related to the original stochastic process. More precisely, a stochastic process that has the same index set , state space , and probability space as another stochastic process is said to be a modification of if for all the following

holds. Two stochastic processes that are modifications of each other have the same finite-dimensional law[158] and they are said to be stochastically equivalent or equivalent.[159]

Instead of modification, the term version is also used,[150][160][161][162] however some authors use the term version when two stochastic processes have the same finite-dimensional distributions, but they may be defined on different probability spaces, so two processes that are modifications of each other, are also versions of each other, in the latter sense, but not the converse.[163][142]

If a continuous-time real-valued stochastic process meets certain moment conditions on its increments, then the Kolmogorov continuity theorem says that there exists a modification of this process that has continuous sample paths with probability one, so the stochastic process has a continuous modification or version.[161][162][164] The theorem can also be generalized to random fields so the index set is -dimensional Euclidean space[165] as well as to stochastic processes with metric spaces as their state spaces.[166]

Indistinguishable

[edit]

Two stochastic processes and defined on the same probability space with the same index set and set space are said be indistinguishable if the following

holds.[142][158] If two and are modifications of each other and are almost surely continuous, then and are indistinguishable.[167]

Separability

[edit]

Separability is a property of a stochastic process based on its index set in relation to the probability measure. The property is assumed so that functionals of stochastic processes or random fields with uncountable index sets can form random variables. For a stochastic process to be separable, in addition to other conditions, its index set must be a separable space,[b] which means that the index set has a dense countable subset.[150][168]

More precisely, a real-valued continuous-time stochastic process on a probability space is separable iff its index set has a dense countable subset and there is a set of probability zero, so , such that for every open set and every closed set , the two events and differ from each other at most on a subset of .[169][170][171] The definition of separability[c] can also be stated for other index sets and state spaces,[174] such as in the case of random fields, where the index set as well as the state space can be -dimensional Euclidean space.[30][150]

The concept of separability of a stochastic process was introduced by Joseph Doob.[168] The underlying idea of separability is to make a countable set of points of the index set determine the properties of the stochastic process.[172] Any stochastic process with a countable index set already meets the separability conditions, so discrete-time stochastic processes are always separable.[175] A theorem by Doob, sometimes known as Doob’s separability theorem, says that any real-valued continuous-time stochastic process has a separable modification.[168][170][176] Versions of this theorem also exist for more general stochastic processes with index sets and state spaces other than the real line.[136]

Independence

[edit]

Two stochastic processes and defined on the same probability space with the same index set are said be independent if for all and for every choice of epochs , the random vectors and are independent.[177]: p. 515 

Uncorrelatedness

[edit]

Two stochastic processes and are called uncorrelated if their cross-covariance is zero for all times.[178]: p. 142  Formally:

.

Independence implies uncorrelatedness

[edit]

If two stochastic processes and are independent, then they are also uncorrelated.[178]: p. 151 

Orthogonality

[edit]

Two stochastic processes and are called orthogonal if their cross-correlation is zero for all times.[178]: p. 142  Formally:

.

Skorokhod space

[edit]

A Skorokhod space, also written as Skorohod space, is a mathematical space of all the functions that are right-continuous with left limits, defined on some interval of the real line such as or , and take values on the real line or on some metric space.[179][180][181] Such functions are known as càdlàg or cadlag functions, based on the acronym of the French phrase continue à droite, limite à gauche.[179][182] A Skorokhod function space, introduced by Anatoliy Skorokhod,[181] is often denoted with the letter ,[179][180][181][182] so the function space is also referred to as space .[179][183][184] The notation of this function space can also include the interval on which all the càdlàg functions are defined, so, for example, denotes the space of càdlàg functions defined on the unit interval .[182][184][185]

Skorokhod function spaces are frequently used in the theory of stochastic processes because it often assumed that the sample functions of continuous-time stochastic processes belong to a Skorokhod space.[181][183] Such spaces contain continuous functions, which correspond to sample functions of the Wiener process. But the space also has functions with discontinuities, which means that the sample functions of stochastic processes with jumps, such as the Poisson process (on the real line), are also members of this space.[184][186]

Regularity

[edit]

In the context of mathematical construction of stochastic processes, the term regularity is used when discussing and assuming certain conditions for a stochastic process to resolve possible construction issues.[187][188] For example, to study stochastic processes with uncountable index sets, it is assumed that the stochastic process adheres to some type of regularity condition such as the sample functions being continuous.[189][190]

Further examples

[edit]

Markov processes and chains

[edit]

Markov processes are stochastic processes, traditionally in discrete or continuous time, that have the Markov property, which means the next value of the Markov process depends on the current value, but it is conditionally independent of the previous values of the stochastic process. In other words, the behavior of the process in the future is stochastically independent of its behavior in the past, given the current state of the process.[191][192]

The Brownian motion process and the Poisson process (in one dimension) are both examples of Markov processes[193] in continuous time, while random walks on the integers and the gambler's ruin problem are examples of Markov processes in discrete time.[194][195]

A Markov chain is a type of Markov process that has either discrete state space or discrete index set (often representing time), but the precise definition of a Markov chain varies.[196] For example, it is common to define a Markov chain as a Markov process in either discrete or continuous time with a countable state space (thus regardless of the nature of time),[197][198][199][200] but it has been also common to define a Markov chain as having discrete time in either countable or continuous state space (thus regardless of the state space).[196] It has been argued that the first definition of a Markov chain, where it has discrete time, now tends to be used, despite the second definition having been used by researchers like Joseph Doob and Kai Lai Chung.[201]

Markov processes form an important class of stochastic processes and have applications in many areas.[39][202] For example, they are the basis for a general stochastic simulation method known as Markov chain Monte Carlo, which is used for simulating random objects with specific probability distributions, and has found application in Bayesian statistics.[203][204]

The concept of the Markov property was originally for stochastic processes in continuous and discrete time, but the property has been adapted for other index sets such as -dimensional Euclidean space, which results in collections of random variables known as Markov random fields.[205][206][207]

Martingale

[edit]

A martingale is a discrete-time or continuous-time stochastic process with the property that, at every instant, given the current value and all the past values of the process, the conditional expectation of every future value is equal to the current value. In discrete time, if this property holds for the next value, then it holds for all future values. The exact mathematical definition of a martingale requires two other conditions coupled with the mathematical concept of a filtration, which is related to the intuition of increasing available information as time passes. Martingales are usually defined to be real-valued,[208][209][155] but they can also be complex-valued[210] or even more general.[211]

A symmetric random walk and a Wiener process (with zero drift) are both examples of martingales, respectively, in discrete and continuous time.[208][209] For a sequence of independent and identically distributed random variables with zero mean, the stochastic process formed from the successive partial sums is a discrete-time martingale.[212] In this aspect, discrete-time martingales generalize the idea of partial sums of independent random variables.[213]

Martingales can also be created from stochastic processes by applying some suitable transformations, which is the case for the homogeneous Poisson process (on the real line) resulting in a martingale called the compensated Poisson process.[209] Martingales can also be built from other martingales.[212] For example, there are martingales based on the martingale the Wiener process, forming continuous-time martingales.[208][214]

Martingales mathematically formalize the idea of a 'fair game' where it is possible form reasonable expectations for payoffs,[215] and they were originally developed to show that it is not possible to gain an 'unfair' advantage in such a game.[216] But now they are used in many areas of probability, which is one of the main reasons for studying them.[155][216][217] Many problems in probability have been solved by finding a martingale in the problem and studying it.[218] Martingales will converge, given some conditions on their moments, so they are often used to derive convergence results, due largely to martingale convergence theorems.[213][219][220]

Martingales have many applications in statistics, but it has been remarked that its use and application are not as widespread as it could be in the field of statistics, particularly statistical inference.[221] They have found applications in areas in probability theory such as queueing theory and Palm calculus[222] and other fields such as economics[223] and finance.[17]

Lévy process

[edit]

Lévy processes are types of stochastic processes that can be considered as generalizations of random walks in continuous time.[49][224] These processes have many applications in fields such as finance, fluid mechanics, physics and biology.[225][226] The main defining characteristics of these processes are their stationarity and independence properties, so they were known as processes with stationary and independent increments. In other words, a stochastic process is a Lévy process if for non-negatives numbers, , the corresponding increments

are all independent of each other, and the distribution of each increment only depends on the difference in time.[49]

A Lévy process can be defined such that its state space is some abstract mathematical space, such as a Banach space, but the processes are often defined so that they take values in Euclidean space. The index set is the non-negative numbers, so , which gives the interpretation of time. Important stochastic processes such as the Wiener process, the homogeneous Poisson process (in one dimension), and subordinators are all Lévy processes.[49][224]

Random field

[edit]

A random field is a collection of random variables indexed by a -dimensional Euclidean space or some manifold. In general, a random field can be considered an example of a stochastic or random process, where the index set is not necessarily a subset of the real line.[30] But there is a convention that an indexed collection of random variables is called a random field when the index has two or more dimensions.[5][28][227] If the specific definition of a stochastic process requires the index set to be a subset of the real line, then the random field can be considered as a generalization of stochastic process.[228]

Point process

[edit]

A point process is a collection of points randomly located on some mathematical space such as the real line, -dimensional Euclidean space, or more abstract spaces. Sometimes the term point process is not preferred, as historically the word process denoted an evolution of some system in time, so a point process is also called a random point field.[229] There are different interpretations of a point process, such a random counting measure or a random set.[230][231] Some authors regard a point process and stochastic process as two different objects such that a point process is a random object that arises from or is associated with a stochastic process,[232][233] though it has been remarked that the difference between point processes and stochastic processes is not clear.[233]

Other authors consider a point process as a stochastic process, where the process is indexed by sets of the underlying space[d] on which it is defined, such as the real line or -dimensional Euclidean space.[236][237] Other stochastic processes such as renewal and counting processes are studied in the theory of point processes.[238][233]

History

[edit]

Early probability theory

[edit]

Probability theory has its origins in games of chance, which have a long history, with some games being played thousands of years ago,[239] but very little analysis on them was done in terms of probability.[240] The year 1654 is often considered the birth of probability theory when French mathematicians Pierre Fermat and Blaise Pascal had a written correspondence on probability, motivated by a gambling problem.[241][242] But there was earlier mathematical work done on the probability of gambling games such as Liber de Ludo Aleae by Gerolamo Cardano, written in the 16th century but posthumously published later in 1663.[243]

After Cardano, Jakob Bernoulli[e] wrote Ars Conjectandi, which is considered a significant event in the history of probability theory. Bernoulli's book was published, also posthumously, in 1713 and inspired many mathematicians to study probability.[245][246] But despite some renowned mathematicians contributing to probability theory, such as Pierre-Simon Laplace, Abraham de Moivre, Carl Gauss, Siméon Poisson and Pafnuty Chebyshev,[247][248] most of the mathematical community[f] did not consider probability theory to be part of mathematics until the 20th century.[247][249][250][251]

Statistical mechanics

[edit]

In the physical sciences, scientists developed in the 19th century the discipline of statistical mechanics, where physical systems, such as containers filled with gases, are regarded or treated mathematically as collections of many moving particles. Although there were attempts to incorporate randomness into statistical physics by some scientists, such as Rudolf Clausius, most of the work had little or no randomness.[252][253] This changed in 1859 when James Clerk Maxwell contributed significantly to the field, more specifically, to the kinetic theory of gases, by presenting work where he modelled the gas particles as moving in random directions at random velocities.[254][255] The kinetic theory of gases and statistical physics continued to be developed in the second half of the 19th century, with work done chiefly by Clausius, Ludwig Boltzmann and Josiah Gibbs, which would later have an influence on Albert Einstein's mathematical model for Brownian movement.[256]

Measure theory and probability theory

[edit]

At the International Congress of Mathematicians in Paris in 1900, David Hilbert presented a list of mathematical problems, where his sixth problem asked for a mathematical treatment of physics and probability involving axioms.[248] Around the start of the 20th century, mathematicians developed measure theory, a branch of mathematics for studying integrals of mathematical functions, where two of the founders were French mathematicians, Henri Lebesgue and Émile Borel. In 1925, another French mathematician Paul Lévy published the first probability book that used ideas from measure theory.[248]

In the 1920s, fundamental contributions to probability theory were made in the Soviet Union by mathematicians such as Sergei Bernstein, Aleksandr Khinchin,[g] and Andrei Kolmogorov.[251] Kolmogorov published in 1929 his first attempt at presenting a mathematical foundation, based on measure theory, for probability theory.[257] In the early 1930s, Khinchin and Kolmogorov set up probability seminars, which were attended by researchers such as Eugene Slutsky and Nikolai Smirnov,[258] and Khinchin gave the first mathematical definition of a stochastic process as a set of random variables indexed by the real line.[63][259][h]

Birth of modern probability theory

[edit]

In 1933, Andrei Kolmogorov published in German, his book on the foundations of probability theory titled Grundbegriffe der Wahrscheinlichkeitsrechnung,[i] where Kolmogorov used measure theory to develop an axiomatic framework for probability theory. The publication of this book is now widely considered to be the birth of modern probability theory, when the theories of probability and stochastic processes became parts of mathematics.[248][251]

After the publication of Kolmogorov's book, further fundamental work on probability theory and stochastic processes was done by Khinchin and Kolmogorov as well as other mathematicians such as Joseph Doob, William Feller, Maurice Fréchet, Paul Lévy, Wolfgang Doeblin, and Harald Cramér.[248][251] Decades later, Cramér referred to the 1930s as the "heroic period of mathematical probability theory".[251] World War II greatly interrupted the development of probability theory, causing, for example, the migration of Feller from Sweden to the United States of America[251] and the death of Doeblin, considered now a pioneer in stochastic processes.[261]

Mathematician Joseph Doob did early work on the theory of stochastic processes, making fundamental contributions, particularly in the theory of martingales.[262][260] His book Stochastic Processes is considered highly influential in the field of probability theory.[263]

Stochastic processes after World War II

[edit]

After World War II, the study of probability theory and stochastic processes gained more attention from mathematicians, with significant contributions made in many areas of probability and mathematics as well as the creation of new areas.[251][264] Starting in the 1940s, Kiyosi Itô published papers developing the field of stochastic calculus, which involves stochastic integrals and stochastic differential equations based on the Wiener or Brownian motion process.[265]

Also starting in the 1940s, connections were made between stochastic processes, particularly martingales, and the mathematical field of potential theory, with early ideas by Shizuo Kakutani and then later work by Joseph Doob.[264] Further work, considered pioneering, was done by Gilbert Hunt in the 1950s, connecting Markov processes and potential theory, which had a significant effect on the theory of Lévy processes and led to more interest in studying Markov processes with methods developed by Itô.[21][266][267]

In 1953, Doob published his book Stochastic processes, which had a strong influence on the theory of stochastic processes and stressed the importance of measure theory in probability.[264] [263] Doob also chiefly developed the theory of martingales, with later substantial contributions by Paul-André Meyer. Earlier work had been carried out by Sergei Bernstein, Paul Lévy and Jean Ville, the latter adopting the term martingale for the stochastic process.[268][269] Methods from the theory of martingales became popular for solving various probability problems. Techniques and theory were developed to study Markov processes and then applied to martingales. Conversely, methods from the theory of martingales were established to treat Markov processes.[264]

Other fields of probability were developed and used to study stochastic processes, with one main approach being the theory of large deviations.[264] The theory has many applications in statistical physics, among other fields, and has core ideas going back to at least the 1930s. Later in the 1960s and 1970s, fundamental work was done by Alexander Wentzell in the Soviet Union and Monroe D. Donsker and Srinivasa Varadhan in the United States of America,[270] which would later result in Varadhan winning the 2007 Abel Prize.[271] In the 1990s and 2000s the theories of Schramm–Loewner evolution[272] and rough paths[142] were introduced and developed to study stochastic processes and other mathematical objects in probability theory, which respectively resulted in Fields Medals being awarded to Wendelin Werner[273] in 2008 and to Martin Hairer in 2014.[274]

The theory of stochastic processes still continues to be a focus of research, with yearly international conferences on the topic of stochastic processes.[45][225]

Discoveries of specific stochastic processes

[edit]

Although Khinchin gave mathematical definitions of stochastic processes in the 1930s,[63][259] specific stochastic processes had already been discovered in different settings, such as the Brownian motion process and the Poisson process.[21][24] Some families of stochastic processes such as point processes or renewal processes have long and complex histories, stretching back centuries.[275]

Bernoulli process

[edit]

The Bernoulli process, which can serve as a mathematical model for flipping a biased coin, is possibly the first stochastic process to have been studied.[81] The process is a sequence of independent Bernoulli trials,[82] which are named after Jacob Bernoulli who used them to study games of chance, including probability problems proposed and studied earlier by Christiaan Huygens.[276] Bernoulli's work, including the Bernoulli process, were published in his book Ars Conjectandi in 1713.[277]

Random walks

[edit]

In 1905, Karl Pearson coined the term random walk while posing a problem describing a random walk on the plane, which was motivated by an application in biology, but such problems involving random walks had already been studied in other fields. Certain gambling problems that were studied centuries earlier can be considered as problems involving random walks.[89][277] For example, the problem known as the Gambler's ruin is based on a simple random walk,[195][278] and is an example of a random walk with absorbing barriers.[241][279] Pascal, Fermat and Huyens all gave numerical solutions to this problem without detailing their methods,[280] and then more detailed solutions were presented by Jakob Bernoulli and Abraham de Moivre.[281]

For random walks in -dimensional integer lattices, George Pólya published, in 1919 and 1921, work where he studied the probability of a symmetric random walk returning to a previous position in the lattice. Pólya showed that a symmetric random walk, which has an equal probability to advance in any direction in the lattice, will return to a previous position in the lattice an infinite number of times with probability one in one and two dimensions, but with probability zero in three or higher dimensions.[282][283]

Wiener process

[edit]

The Wiener process or Brownian motion process has its origins in different fields including statistics, finance and physics.[21] In 1880, Danish astronomer Thorvald Thiele wrote a paper on the method of least squares, where he used the process to study the errors of a model in time-series analysis.[284][285][286] The work is now considered as an early discovery of the statistical method known as Kalman filtering, but the work was largely overlooked. It is thought that the ideas in Thiele's paper were too advanced to have been understood by the broader mathematical and statistical community at the time.[286]

Norbert Wiener gave the first mathematical proof of the existence of the Wiener process. This mathematical object had appeared previously in the work of Thorvald Thiele, Louis Bachelier, and Albert Einstein.[21]

The French mathematician Louis Bachelier used a Wiener process in his 1900 thesis[287][288] in order to model price changes on the Paris Bourse, a stock exchange,[289] without knowing the work of Thiele.[21] It has been speculated that Bachelier drew ideas from the random walk model of Jules Regnault, but Bachelier did not cite him,[290] and Bachelier's thesis is now considered pioneering in the field of financial mathematics.[289][290]

It is commonly thought that Bachelier's work gained little attention and was forgotten for decades until it was rediscovered in the 1950s by the Leonard Savage, and then become more popular after Bachelier's thesis was translated into English in 1964. But the work was never forgotten in the mathematical community, as Bachelier published a book in 1912 detailing his ideas,[290] which was cited by mathematicians including Doob, Feller[290] and Kolmogorov.[21] The book continued to be cited, but then starting in the 1960s, the original thesis by Bachelier began to be cited more than his book when economists started citing Bachelier's work.[290]

In 1905, Albert Einstein published a paper where he studied the physical observation of Brownian motion or movement to explain the seemingly random movements of particles in liquids by using ideas from the kinetic theory of gases. Einstein derived a differential equation, known as a diffusion equation, for describing the probability of finding a particle in a certain region of space. Shortly after Einstein's first paper on Brownian movement, Marian Smoluchowski published work where he cited Einstein, but wrote that he had independently derived the equivalent results by using a different method.[291]

Einstein's work, as well as experimental results obtained by Jean Perrin, later inspired Norbert Wiener in the 1920s[292] to use a type of measure theory, developed by Percy Daniell, and Fourier analysis to prove the existence of the Wiener process as a mathematical object.[21]

Poisson process

[edit]

The Poisson process is named after Siméon Poisson, due to its definition involving the Poisson distribution, but Poisson never studied the process.[22][293] There are a number of claims for early uses or discoveries of the Poisson process.[22][24] At the beginning of the 20th century, the Poisson process would arise independently in different situations.[22][24] In Sweden 1903, Filip Lundberg published a thesis containing work, now considered fundamental and pioneering, where he proposed to model insurance claims with a homogeneous Poisson process.[294][295]

Another discovery occurred in Denmark in 1909 when A.K. Erlang derived the Poisson distribution when developing a mathematical model for the number of incoming phone calls in a finite time interval. Erlang was not at the time aware of Poisson's earlier work and assumed that the number phone calls arriving in each interval of time were independent to each other. He then found the limiting case, which is effectively recasting the Poisson distribution as a limit of the binomial distribution.[22]

In 1910, Ernest Rutherford and Hans Geiger published experimental results on counting alpha particles. Motivated by their work, Harry Bateman studied the counting problem and derived Poisson probabilities as a solution to a family of differential equations, resulting in the independent discovery of the Poisson process.[22] After this time there were many studies and applications of the Poisson process, but its early history is complicated, which has been explained by the various applications of the process in numerous fields by biologists, ecologists, engineers and various physical scientists.[22]

Markov processes

[edit]

Markov processes and Markov chains are named after Andrey Markov who studied Markov chains in the early 20th century. Markov was interested in studying an extension of independent random sequences. In his first paper on Markov chains, published in 1906, Markov showed that under certain conditions the average outcomes of the Markov chain would converge to a fixed vector of values, so proving a weak law of large numbers without the independence assumption,[296][297][298] which had been commonly regarded as a requirement for such mathematical laws to hold.[298] Markov later used Markov chains to study the distribution of vowels in Eugene Onegin, written by Alexander Pushkin, and proved a central limit theorem for such chains.

In 1912, Poincaré studied Markov chains on finite groups with an aim to study card shuffling. Other early uses of Markov chains include a diffusion model, introduced by Paul and Tatyana Ehrenfest in 1907, and a branching process, introduced by Francis Galton and Henry William Watson in 1873, preceding the work of Markov.[296][297] After the work of Galton and Watson, it was later revealed that their branching process had been independently discovered and studied around three decades earlier by Irénée-Jules Bienaymé.[299] Starting in 1928, Maurice Fréchet became interested in Markov chains, eventually resulting in him publishing in 1938 a detailed study on Markov chains.[296][300]

Andrei Kolmogorov developed in a 1931 paper a large part of the early theory of continuous-time Markov processes.[251][257] Kolmogorov was partly inspired by Louis Bachelier's 1900 work on fluctuations in the stock market as well as Norbert Wiener's work on Einstein's model of Brownian movement.[257][301] He introduced and studied a particular set of Markov processes known as diffusion processes, where he derived a set of differential equations describing the processes.[257][302] Independent of Kolmogorov's work, Sydney Chapman derived in a 1928 paper an equation, now called the Chapman–Kolmogorov equation, in a less mathematically rigorous way than Kolmogorov, while studying Brownian movement.[303] The differential equations are now called the Kolmogorov equations[304] or the Kolmogorov–Chapman equations.[305] Other mathematicians who contributed significantly to the foundations of Markov processes include William Feller, starting in the 1930s, and then later Eugene Dynkin, starting in the 1950s.[251]

Lévy processes

[edit]

Lévy processes such as the Wiener process and the Poisson process (on the real line) are named after Paul Lévy who started studying them in the 1930s,[225] but they have connections to infinitely divisible distributions going back to the 1920s.[224] In a 1932 paper, Kolmogorov derived a characteristic function for random variables associated with Lévy processes. This result was later derived under more general conditions by Lévy in 1934, and then Khinchin independently gave an alternative form for this characteristic function in 1937.[251][306] In addition to Lévy, Khinchin and Kolomogrov, early fundamental contributions to the theory of Lévy processes were made by Bruno de Finetti and Kiyosi Itô.[224]

Mathematical construction

[edit]

In mathematics, constructions of mathematical objects are needed, which is also the case for stochastic processes, to prove that they exist mathematically.[57] There are two main approaches for constructing a stochastic process. One approach involves considering a measurable space of functions, defining a suitable measurable mapping from a probability space to this measurable space of functions, and then deriving the corresponding finite-dimensional distributions.[307]

Another approach involves defining a collection of random variables to have specific finite-dimensional distributions, and then using Kolmogorov's existence theorem[j] to prove a corresponding stochastic process exists.[57][307] This theorem, which is an existence theorem for measures on infinite product spaces,[311] says that if any finite-dimensional distributions satisfy two conditions, known as consistency conditions, then there exists a stochastic process with those finite-dimensional distributions.[57]

Construction issues

[edit]

When constructing continuous-time stochastic processes certain mathematical difficulties arise, due to the uncountable index sets, which do not occur with discrete-time processes.[58][59] One problem is that it is possible to have more than one stochastic process with the same finite-dimensional distributions. For example, both the left-continuous modification and the right-continuous modification of a Poisson process have the same finite-dimensional distributions.[312] This means that the distribution of the stochastic process does not, necessarily, specify uniquely the properties of the sample functions of the stochastic process.[307][313]

Another problem is that functionals of continuous-time process that rely upon an uncountable number of points of the index set may not be measurable, so the probabilities of certain events may not be well-defined.[168] For example, the supremum of a stochastic process or random field is not necessarily a well-defined random variable.[30][59] For a continuous-time stochastic process , other characteristics that depend on an uncountable number of points of the index set include:[168]

  • a sample function of a stochastic process is a continuous function of ;
  • a sample function of a stochastic process is a bounded function of ; and
  • a sample function of a stochastic process is an increasing function of .

where the symbol can be read "a member of the set", as in a member of the set .

To overcome the two difficulties described above, i.e., "more than one..." and "functionals of...", different assumptions and approaches are possible.[69]

Resolving construction issues

[edit]

One approach for avoiding mathematical construction issues of stochastic processes, proposed by Joseph Doob, is to assume that the stochastic process is separable.[314] Separability ensures that infinite-dimensional distributions determine the properties of sample functions by requiring that sample functions are essentially determined by their values on a dense countable set of points in the index set.[315] Furthermore, if a stochastic process is separable, then functionals of an uncountable number of points of the index set are measurable and their probabilities can be studied.[168][315]

Another approach is possible, originally developed by Anatoliy Skorokhod and Andrei Kolmogorov,[316] for a continuous-time stochastic process with any metric space as its state space. For the construction of such a stochastic process, it is assumed that the sample functions of the stochastic process belong to some suitable function space, which is usually the Skorokhod space consisting of all right-continuous functions with left limits. This approach is now more used than the separability assumption,[69][262] but such a stochastic process based on this approach will be automatically separable.[317]

Although less used, the separability assumption is considered more general because every stochastic process has a separable version.[262] It is also used when it is not possible to construct a stochastic process in a Skorokhod space.[173] For example, separability is assumed when constructing and studying random fields, where the collection of random variables is now indexed by sets other than the real line such as -dimensional Euclidean space.[30][318]

Application

[edit]

Applications in Finance

[edit]

Black-Scholes Model

[edit]

One of the most famous applications of stochastic processes in finance is the Black-Scholes model for option pricing. Developed by Fischer Black, Myron Scholes, and Robert Solow, this model uses Geometric Brownian motion, a specific type of stochastic process, to describe the dynamics of asset prices.[319][320] The model assumes that the price of a stock follows a continuous-time stochastic process and provides a closed-form solution for pricing European-style options. The Black-Scholes formula has had a profound impact on financial markets, forming the basis for much of modern options trading.

The key assumption of the Black-Scholes model is that the price of a financial asset, such as a stock, follows a log-normal distribution, with its continuous returns following a normal distribution. Although the model has limitations, such as the assumption of constant volatility, it remains widely used due to its simplicity and practical relevance.

Stochastic Volatility Models

[edit]

Another significant application of stochastic processes in finance is in stochastic volatility models, which aim to capture the time-varying nature of market volatility. The Heston model[321] is a popular example, allowing for the volatility of asset prices to follow its own stochastic process. Unlike the Black-Scholes model, which assumes constant volatility, stochastic volatility models provide a more flexible framework for modeling market dynamics, particularly during periods of high uncertainty or market stress.

Applications in Biology

[edit]

Population Dynamics

[edit]

One of the primary applications of stochastic processes in biology is in population dynamics. In contrast to deterministic models, which assume that populations change in predictable ways, stochastic models account for the inherent randomness in births, deaths, and migration. The birth-death process,[322] a simple stochastic model, describes how populations fluctuate over time due to random births and deaths. These models are particularly important when dealing with small populations, where random events can have large impacts, such as in the case of endangered species or small microbial populations.

Another example is the branching process,[322] which models the growth of a population where each individual reproduces independently. The branching process is often used to describe population extinction or explosion, particularly in epidemiology, where it can model the spread of infectious diseases within a population.

Applications in Computer Science

[edit]

Randomized Algorithms

[edit]

Stochastic processes play a critical role in computer science, particularly in the analysis and development of randomized algorithms. These algorithms utilize random inputs to simplify problem-solving or enhance performance in complex computational tasks. For instance, Markov chains are widely used in probabilistic algorithms for optimization and sampling tasks, such as those employed in search engines like Google's PageRank.[323] These methods balance computational efficiency with accuracy, making them invaluable for handling large datasets. Randomized algorithms are also extensively applied in areas such as cryptography, large-scale simulations, and artificial intelligence, where uncertainty must be managed effectively.[323]

Queuing Theory

[edit]

Another significant application of stochastic processes in computer science is in queuing theory, which models the random arrival and service of tasks in a system.[324] This is particularly relevant in network traffic analysis and server management. For instance, queuing models help predict delays, manage resource allocation, and optimize throughput in web servers and communication networks. The flexibility of stochastic models allows researchers to simulate and improve the performance of high-traffic environments. For example, queueing theory is crucial for designing efficient data centers and cloud computing infrastructures.[325]

See also

[edit]

Notes

[edit]

References

[edit]

Further reading

[edit]
[edit]
Revisions and contributorsEdit on WikipediaRead on Wikipedia
from Grokipedia
A stochastic process is a mathematical object that models a sequence of random variables evolving over time or another index set, providing a framework to describe systems subject to uncertainty and randomness.[1] Formally, it is defined as a family of random variables {Xt:tT}\{X_t : t \in T\}, where TT is the index set (often time, either discrete like integers or continuous like reals), and each XtX_t represents the state of the system at index tt.[2] This structure captures the probabilistic evolution of phenomena where outcomes are not deterministic but governed by probability distributions.[3] Stochastic processes are classified based on several criteria, including the nature of the index set and the state space, leading to discrete-time processes (where TT is countable) and continuous-time processes (where TT is uncountable).[4] Key types include Markov processes, which depend only on the current state rather than the full history; random walks, modeling step-by-step random movements; Poisson processes, describing event occurrences at constant average rates; and Brownian motion, a continuous-time process with independent, normally distributed increments.[5] Additional categories encompass Gaussian processes (with jointly normal marginal distributions), processes with independent increments, and stationary processes (where statistical properties remain invariant over time).[6] These classifications enable tailored modeling of diverse random phenomena. The development of stochastic processes traces back to the late 19th and early 20th centuries, with foundational work on Brownian motion by Louis Bachelier in 1900 for financial modeling and Albert Einstein in 1905 for physical diffusion.[7] With foundational contributions including Norbert Wiener's construction of the Wiener process in 1923 and Andrey Kolmogorov's axiomatic probability theory in 1933 providing a rigorous measure-theoretic foundation, these advancements formalized continuous processes.[8] This historical progression transformed stochastic processes from ad hoc models into a cornerstone of modern probability theory. Applications of stochastic processes span numerous fields, including finance for pricing derivatives and risk assessment via models like geometric Brownian motion; physics and engineering for simulating particle diffusion, queueing systems, and signal processing; biology for population dynamics and genetic drift; and computer science for algorithms in machine learning and network analysis. In operations research, renewal and branching processes optimize resource allocation and reliability engineering.[9] These models are essential for handling real-world uncertainty, enabling predictions and simulations where deterministic approaches fall short.

Introduction and Fundamentals

Overview and Basic Definition

A stochastic process is a mathematical model that describes a sequence of random variables evolving over time or space, capturing the inherent uncertainty in systems such as fluctuating stock prices or the erratic motion of particles in a fluid.[10] These processes provide a framework for analyzing phenomena where outcomes are probabilistic rather than deterministic, allowing researchers to quantify risks, predict trends, and simulate behaviors in fields ranging from finance to physics.[11] The term "stochastic" originates from the Greek word stokhastikos, meaning "skillful in aiming" or "pertaining to guesswork," reflecting its roots in conjecture and probabilistic reasoning.[12] This etymology underscores the early association of such models with uncertainty and estimation, evolving from ancient notions of chance to modern rigorous theory.[13] At its foundation, a stochastic process is defined within a probability space (Ω,F,P)(\Omega, \mathcal{F}, P), where Ω\Omega is the sample space, F\mathcal{F} is a σ\sigma-algebra of events, and PP is a probability measure; the process itself is a family of random variables X=(Xt)tTX = (X_t)_{t \in T}, with each Xt:ΩSX_t: \Omega \to S mapping outcomes to a state space SS for indices tt in an index set TT.[11] Early applications emerged in the 18th century, notably in Jacob Bernoulli's 1713 work Ars Conjectandi, which explored sequences of coin tosses to establish foundational principles like the law of large numbers, initially in the context of gambling but with implications for broader probabilistic modeling.[14]

Classifications by Index Set and State Space

Stochastic processes are classified according to the structure of their index set, which parameterizes the evolution of the process (often time or space), and their state space, which comprises the possible values the process can take. These classifications determine the appropriate mathematical tools, from basic probability for simpler cases to advanced measure theory for more complex ones.[15][16] The index set can be discrete or continuous. A discrete index set consists of a countable collection of points, such as the integers N0={0,1,2,}\mathbb{N}_0 = \{0, 1, 2, \dots\}, modeling processes that update at specific intervals like daily observations. This structure yields countable sample paths, enabling straightforward analysis via recursion and finite computations.[17][15] In contrast, a continuous index set forms an uncountable set, such as the non-negative reals [0,)[0, \infty), suitable for phenomena evolving without discrete jumps, like physical motion. Here, sample paths are uncountable functions, necessitating tools from functional analysis and stochastic integration for proper definition and study.[17][16] The state space is similarly categorized as discrete or continuous. A discrete state space is countable, either finite (e.g., a set of categories) or countably infinite (e.g., non-negative integers for counts), facilitating exact probability calculations through summation and matrix representations. Continuous state spaces are uncountable, often intervals on the real line R\mathbb{R}, as in measurements of position or value, requiring probability densities and integrals for marginal distributions.[15][17] Integrating these dimensions produces hybrid categories: discrete-time discrete-state processes, such as those analyzed via transition matrices; discrete-time continuous-state processes; continuous-time discrete-state processes, like counting arrivals; and continuous-time continuous-state processes, involving diffusion approximations. These combinations influence modeling choices, with discrete variants offering computational ease for simulations and approximations, while continuous ones capture realistic dynamics in fields like finance and physics but demand rigorous probabilistic frameworks.[15][16]

Notation and Terminology

In stochastic processes, standard notation denotes a process as $ X = (X_t)_{t \in T} $, where $ {X_t : t \in T} $ is a family of random variables indexed by the set $ T $, the index set, taking values in the state space $ E $, and defined on the underlying probability space $ (\Omega, \mathcal{F}, P) $, with $ \Omega $ the sample space, $ \mathcal{F} $ the sigma-algebra, and $ P $ the probability measure.[11] The term stochastic process refers to the abstract collection of these random variables $ X_t $, each representing the state at index $ t $. A realization or sample path of the process is a specific outcome $ \omega \in \Omega $, yielding the deterministic function $ t \mapsto X_t(\omega) $ from $ T $ to $ E $, which traces the evolution of the process for that particular sample. The law of the process describes its probabilistic structure, fully determined by the finite-dimensional distributions of the family $ (X_{t_1}, \dots, X_{t_n}) $ for any finite $ n $ and $ t_1, \dots, t_n \in T $.[11][18] Common abbreviations include i.i.d. for independent and identically distributed random variables, meaning the variables are mutually independent and share the same probability distribution. Another standard term is CDF for cumulative distribution function, which for a random variable $ X $ is the function $ F_X(x) = P(X \leq x) $, providing the probability that $ X $ does not exceed $ x $.[19][20] For path regularity, a key convention in continuous-time processes is the assumption of right-continuous paths, where $ \lim_{s \downarrow t} X_s = X_t $ for each $ t \in T $. More generally, processes with possible jumps, such as counting processes, are often taken to have càdlàg paths—right-continuous with left limits—derived from the French phrase continu à droite, limite à gauche, ensuring $ \lim_{s \downarrow t} X_s = X_t $ and $ \lim_{s \uparrow t} X_s $ exists for all $ t $.[21]

Core Examples

Bernoulli Process

The Bernoulli process is a fundamental discrete-time stochastic process consisting of an infinite sequence of independent and identically distributed (i.i.d.) Bernoulli random variables $ {X_n : n = 1, 2, \dots } $, where each $ X_n $ takes the value 1 with probability $ p $ (representing a "success") and 0 with probability $ 1-p $ (representing a "failure"), with $ 0 < p < 1 $.[22][23][24] This process models sequences of binary trials, such as repeated coin flips or independent detections in a signal processing context, where the outcome of each trial does not influence the others.[22][24] A key feature of the Bernoulli process is the partial sum process $ S_n = \sum_{k=1}^n X_k $, which counts the number of successes up to time $ n $ and follows a binomial distribution with parameters $ n $ and $ p $.[23][22] The expected value of this sum is $ \mathbb{E}[S_n] = np $, reflecting the average number of successes over $ n $ trials, while the variance is $ \mathrm{Var}(S_n) = np(1-p) $, capturing the variability due to the binary nature of the outcomes.[23][22] The process exhibits several important properties that underscore its simplicity and utility. The increments $ X_{n+1}, X_{n+2}, \dots $ are independent of the past $ {X_1, \dots, X_n} $, ensuring that future trials remain unaffected by prior results—a property known as memorylessness.[22][24] Additionally, it is stationary, meaning the joint distribution of $ {X_{m+1}, \dots, X_{m+k}} $ is identical to that of $ {X_1, \dots, X_k} $ for any $ m $, due to the constant success probability $ p $.[23] This direct link to the binomial distribution for the partial sums makes the Bernoulli process a cornerstone for understanding counting processes in probability.[23][22] As a basic model of independent binary events, the Bernoulli process serves as the foundation for more elaborate stochastic models, such as the simple random walk, where the partial sums track cumulative positions.[24]

Random Walk

The simple symmetric random walk is a discrete-time stochastic process that models the position of a particle taking successive random steps of equal length on the integer lattice, serving as a foundational example that illustrates accumulation of independent random increments and connects to asymptotic behaviors like the central limit theorem. Formally, the position at step nn, denoted SnS_n, is given by the partial sum
Sn=k=1nYk, S_n = \sum_{k=1}^n Y_k,
where S0=0S_0 = 0 and each increment YkY_k is an independent random variable taking value +1+1 or 1-1 with probability 1/21/2 each.[25][26] The increments {Yk}\{Y_k\} are independent and identically distributed (stationary), with mean zero and variance one, implying that SnS_n has mean zero and variance nn.[27][28] In one dimension, the probability of returning to the origin after 2n2n steps is (2nn)(1/2)2n\binom{2n}{n} (1/2)^{2n}, and the infinite sum of these probabilities over nn diverges, indicating recurrence.[29] This process is recurrent in one and two dimensions—returning to the starting point with probability one—but transient in three or more dimensions, where the return probability is less than one, as proven by Pólya's theorem.[30][31] Asymptotically, a properly scaled and centered version of the simple symmetric random walk converges in distribution to a standard Brownian motion, bridging discrete and continuous stochastic models.[32]

Poisson Process

The Poisson process is a fundamental continuous-time stochastic process used to model the occurrence of rare events, such as arrivals or incidents, over time. It is defined as a counting process $ {N(t) : t \geq 0} $, where $ N(t) $ represents the number of events that have occurred by time $ t $, starting with $ N(0) = 0 $. The process has independent increments, meaning that the number of events in disjoint time intervals are independent random variables, and stationary increments, meaning that the distribution of the increment $ N(t + s) - N(t) $ depends only on the length $ s $ of the interval. For small $ h > 0 $, the probability of exactly $ k $ events occurring in a short interval $ (t, t + h] $ satisfies $ P(N(t + h) - N(t) = k) \approx \frac{(\lambda h)^k e^{-\lambda h}}{k!} $, where $ \lambda > 0 $ is the constant rate parameter, along with $ P(N(t + h) - N(t) \geq 2) = o(h) $ as $ h \to 0 $.[33][34] A key property of the Poisson process is that the number of events in any fixed interval $ (0, t] $, denoted $ N(t) $, follows a Poisson distribution with parameter $ \lambda t $, so $ N(t) \sim \mathrm{Pois}(\lambda t) $ and $ P(N(t) = n) = \frac{(\lambda t)^n e^{-\lambda t}}{n!} $ for $ n = 0, 1, 2, \dots $. The interarrival times between successive events are independent and exponentially distributed with rate $ \lambda $, meaning the waiting time until the next event has density $ f(x) = \lambda e^{-\lambda x} $ for $ x \geq 0 $. This exponential distribution implies the memoryless property: the distribution of the remaining time until the next event does not depend on how much time has already elapsed. The process is homogeneous, with constant intensity $ \lambda $, and the expected number of events by time $ t $ is $ E[N(t)] = \lambda t $, reflecting a linear growth rate in expectation.[33] The Poisson process exhibits useful superposition and thinning properties that facilitate modeling complex systems from simpler components. Superposition states that the merger of two independent Poisson processes with rates $ \lambda_1 $ and $ \lambda_2 $ results in another Poisson process with rate $ \lambda_1 + \lambda_2 $; this extends to any finite number of independent processes. Thinning, conversely, involves independently classifying each event of a Poisson process with rate $ \lambda $ into types with probabilities $ p $ and $ 1 - p $, yielding two independent Poisson processes with rates $ \lambda p $ and $ \lambda (1 - p) $. These properties underscore the process's role as a building block for more general point processes, including its classification as a continuous-time Lévy process.[33][34]

Wiener Process

The Wiener process, also known as standard Brownian motion, serves as the canonical example of a continuous-time stochastic process with continuous sample paths and Gaussian marginal distributions. It models the random motion of particles suspended in a fluid, as observed in physical phenomena like diffusion, and forms the foundation for many advanced stochastic models in mathematics, physics, and finance.[35] Formally, a Wiener process W={W(t):t0}W = \{W(t) : t \geq 0\} is defined on a probability space (Ω,F,P)(\Omega, \mathcal{F}, P) as a stochastic process satisfying the following properties: W(0)=0W(0) = 0 almost surely; the increments W(t)W(s)W(t) - W(s) for t>s0t > s \geq 0 are independent and normally distributed as W(t)W(s)N(0,ts)W(t) - W(s) \sim \mathcal{N}(0, t - s), meaning the process has independent stationary increments. These conditions ensure that the process is a Lévy process with Gaussian increments, distinguishing it from discrete-time processes like the random walk.[35][36] Key properties of the Wiener process include the almost sure continuity of its sample paths, meaning that with probability 1, the trajectory tW(t,ω)t \mapsto W(t, \omega) is continuous for almost all outcomes ωΩ\omega \in \Omega. The covariance function is given by Cov(W(s),W(t))=min(s,t)\operatorname{Cov}(W(s), W(t)) = \min(s, t) for s,t0s, t \geq 0, which captures the shared randomness up to the earlier time. Additionally, the quadratic variation process satisfies Wt=t\langle W \rangle_t = t almost surely, quantifying the accumulated squared increments over [0,t][0, t]. The process exhibits self-similarity, with the scaling property W(ct)=dcW(t)W(ct) \stackrel{d}{=} \sqrt{c} W(t) for any c>0c > 0, reflecting its fractal-like structure at different time scales.[35][36] Historically, the Wiener process is named after Norbert Wiener, who provided a rigorous mathematical construction in his 1923 paper, proving the existence of such a process with continuous paths. However, its conceptual roots trace back to Albert Einstein's 1905 analysis of Brownian motion, where he derived the diffusion equation and related the mean squared displacement of particles to time via $ \mathbb{E}[(X_t - X_0)^2] = 2Dt $, laying the groundwork for the variance structure of the increments.[37]

Formal Definitions

Index Set and State Space

A stochastic process is defined on an underlying probability space (Ω,F,P)(\Omega, \mathcal{F}, P), where Ω\Omega is the sample space, F\mathcal{F} is a σ\sigma-algebra, and PP is a probability measure. The structural foundation of the process rests on two key components: the index set TT and the state space EE. The index set TT is a partially ordered set (poset), which provides the parameter space over which the process evolves; in general formulations, TT may not be totally ordered, allowing for multiparameter or set-indexed processes, though standard cases assume a total order such as the countable set N\mathbb{N} for discrete-time processes or the interval [0,)[0, \infty) for continuous-time ones.[38] To equip TT with a measurable structure, it is typically endowed with the order topology, generating the order σ\sigma-algebra T\mathcal{T} consisting of sets whose membership depends on the ordering relations in TT./02%3A_Probability_Spaces/2.10%3A_Stochastic_Processes) The state space EE is a measurable space (E,E)(E, \mathcal{E}), where E\mathcal{E} is a σ\sigma-algebra on the set EE that specifies the observable events or outcomes the process can take. In many rigorous treatments, EE is chosen to be a Polish space—a separable and completely metrizable topological space—such as Rd\mathbb{R}^d equipped with its Borel σ\sigma-algebra, to guarantee desirable properties like the existence of regular conditional distributions and tightness for weak convergence.[39] This choice ensures that the space supports a rich theory of measurability without pathological sets, facilitating the study of path properties and limits in stochastic analysis.[18] Formally, the stochastic process XX is a function X:T×ΩEX: T \times \Omega \to E that assigns to each pair (t,ω)T×Ω(t, \omega) \in T \times \Omega a state X(t,ω)EX(t, \omega) \in E. For XX to be a valid stochastic process, it must be measurable with respect to the product σ\sigma-algebra TF\mathcal{T} \otimes \mathcal{F} on T×ΩT \times \Omega and E\mathcal{E} on EE; this joint measurability implies that for each fixed tTt \in T, the section Xt:ΩEX_t: \Omega \to E defined by Xt(ω)=X(t,ω)X_t(\omega) = X(t, \omega) is F/E\mathcal{F}/\mathcal{E}-measurable, making XtX_t a random variable./02%3A_Probability_Spaces/2.10%3A_Stochastic_Processes) Equivalently, XX can be viewed as a random element in the space of functions ETE^T, where ETE^T is endowed with the product σ\sigma-algebra generated by the cylinder sets.[40] This joint measurability requirement ensures compatibility across the index set, allowing the process to be consistently defined and analyzed through its finite-dimensional distributions while avoiding inconsistencies arising from non-measurable pathologies. Without it, the process might not integrate well with the probability measure PP, potentially undermining probabilistic interpretations.[41] In practice, for totally ordered TT and Polish EE, this structure supports the Kolmogorov extension theorem, which constructs the process from consistent finite-dimensional distributions.[39]

Sample Paths and Realizations

A sample path of a stochastic process {Xt:tT}\{X_t : t \in T\} defined on a probability space (Ω,F,P)(\Omega, \mathcal{F}, P) with index set TT and state space EE is the function X(,ω):TEX(\cdot, \omega): T \to E obtained by fixing an outcome ωΩ\omega \in \Omega and mapping each tTt \in T to Xt(ω)EX_t(\omega) \in E.[42] This realization traces the evolution of the process for that particular ω\omega, akin to observing a single trajectory through the state space over the index set.[43] Realizations of stochastic processes often exhibit specific properties almost surely, meaning with probability 1 under the measure PP. For instance, the Wiener process, also known as Brownian motion, has sample paths that are almost surely continuous, ensuring that the function W(,ω):[0,)RW(\cdot, \omega): [0, \infty) \to \mathbb{R} is continuous for almost all ωΩ\omega \in \Omega. This almost sure continuity is a fundamental regularity condition for the Wiener process, distinguishing it from processes with discontinuous paths.[44] The collection of all possible sample paths forms the path space, typically denoted as ETE^T, which is the set of all functions from TT to EE. To define a measurable structure on this space, one equips ETE^T with the cylinder σ\sigma-algebra, generated by sets of the form {xET:(xt1,,xtn)B}\{\mathbf{x} \in E^T : (x_{t_1}, \dots, x_{t_n}) \in B\} for finite nn, indices t1,,tnTt_1, \dots, t_n \in T, and Borel sets BEnB \subseteq E^n.[45] For processes with continuous paths, such as the Wiener process, the path space is often restricted to the subspace C[0,)C[0, \infty) of continuous functions on [0,)[0, \infty), equipped with the cylinder σ\sigma-algebra induced from the Borel σ\sigma-algebra on the uniform topology.[18] Two stochastic processes are versions of each other if they possess the same finite-dimensional distributions, yet their sample paths may differ on sets of positive probability.[46] This distinction allows for processes that are probabilistically equivalent in marginals and joints but realized differently as path functions, such as a discontinuous version versus a continuous modification of the same underlying law.[47]

Finite-Dimensional Distributions

The finite-dimensional distributions (f.d.d.) of a stochastic process {Xt}tT\{X_t\}_{t \in T} taking values in a state space EE consist of the marginal probability laws of the random vectors (Xt1,,Xtn)(X_{t_1}, \dots, X_{t_n}) for every finite collection of distinct indices t1<<tnt_1 < \dots < t_n in the index set TT and every nNn \in \mathbb{N}, defined on the product space EnE^n. These distributions fully specify the law of the process on the cylinder σ\sigma-algebra generated by the coordinate projections, providing a complete probabilistic description without reference to path properties.[48] For such a family of distributions to correspond to an actual stochastic process, they must satisfy consistency conditions: specifically, for any n<mn < m and indices s1<<sms_1 < \dots < s_m in TT, the distribution of (Xsi1,,Xsin)(X_{s_{i_1}}, \dots, X_{s_{i_n}}) must equal the nn-dimensional marginal of the mm-dimensional distribution of (Xs1,,Xsm)(X_{s_1}, \dots, X_{s_m}), where i1<<ini_1 < \dots < i_n are any increasing subsequence. The Kolmogorov extension theorem asserts that if the state space EE is a Polish space (complete separable metric space) and the family of finite-dimensional distributions is consistent in this sense, then there exists a unique probability measure on the product space ETE^T (equipped with the product σ\sigma-algebra) such that the induced distributions on finite-dimensional projections match the given family. This construction ensures the existence of the process as a measurable function from a probability space to ETE^T. The marginal and joint probabilities of the process are directly determined by its finite-dimensional distributions. For instance, the joint cumulative distribution function at points t1<<tnTt_1 < \dots < t_n \in T and x1,,xnEx_1, \dots, x_n \in E is given by
Ft1,,tn(x1,,xn)=P(Xt1x1,,Xtnxn), F_{t_1, \dots, t_n}(x_1, \dots, x_n) = P(X_{t_1} \leq x_1, \dots, X_{t_n} \leq x_n),
which specifies the f.d.d. measure on EnE^n. Similarly, one-dimensional marginals yield the laws P(Xt)P(X_t \in \cdot) for each tTt \in T.[43] Two stochastic processes are equal in law (i.e., have the same distribution as random elements of ETE^T) if and only if their finite-dimensional distributions coincide for all finite sets of times and all nn. This weak specification via f.d.d. forms the minimal data required to determine the probabilistic structure of the process, enabling convergence in distribution to be checked through convergence of these finite-dimensional laws (under additional tightness conditions for path space topologies).[48]

Increments and Stationarity

In stochastic processes, the increment over an interval (s,t](s, t] with t>st > s is defined as ΔX(s,t)=XtXs\Delta X(s,t) = X_t - X_s, representing the change in the process value during that period.[49] A key property is the independence of increments: for disjoint intervals, the increments ΔX(si,ti)\Delta X(s_i, t_i) are independent random variables, which underpins the behavior of many processes like Lévy processes.[3] This independence can be characterized through the finite-dimensional distributions of the process, where the joint law of increments over non-overlapping intervals factors into marginals.[50] Stationarity in stochastic processes refers to the invariance of statistical properties under time shifts. Strict stationarity requires that the joint distribution of {Xt1+h,,Xtk+h}\{X_{t_1 + h}, \dots, X_{t_k + h}\} equals that of {Xt1,,Xtk}\{X_{t_1}, \dots, X_{t_k}\} for any kk, times t1<<tkt_1 < \dots < t_k, and shift h>0h > 0.[51] In contrast, weak (or wide-sense) stationarity is a milder condition, demanding a constant mean E[Xt]=μ\mathbb{E}[X_t] = \mu for all tt and an autocovariance function Cov(Xt,Xt+τ)\text{Cov}(X_t, X_{t+\tau}) that depends only on the lag τ\tau, assuming finite second moments exist.[52] Strict stationarity implies weak stationarity when moments are finite, but the converse does not hold.[51] For increments specifically, stationary increments mean the distribution of ΔX(s,t)=XtXs\Delta X(s,t) = X_t - X_s depends solely on the length tst - s, or equivalently, the law of Xt+hXtX_{t+h} - X_t is independent of tt for fixed h>0h > 0:
Xt+hXt=dXhX0 X_{t+h} - X_t \stackrel{d}{=} X_h - X_0
for all t0t \geq 0.[53] Processes with both stationary and independent increments, such as the Poisson process—where increments follow a Poisson distribution with parameter λ(ts)\lambda (t - s)—and the Wiener process—where increments are normally distributed with mean 0 and variance tst - s—exemplify this property and form the basis for Lévy processes.[54][55][56] Ergodicity extends stationarity by ensuring that time averages along a single sample path converge almost surely to ensemble (expectation) averages, allowing inference of global statistics from long realizations of stationary processes.[57] This property holds for many ergodic stationary processes but requires additional mixing conditions beyond mere stationarity.[58]

Key Properties and Structures

Filtrations and Adaptability

In stochastic processes, a filtration provides a mathematical framework for modeling the evolution of available information over time. Formally, given a probability space (Ω,F,P)(\Omega, \mathcal{F}, P) and an index set TT (typically [0,)[0, \infty) or N\mathbb{N}), a filtration is a family of sub-σ\sigma-algebras {Ft}tT\{\mathcal{F}_t\}_{t \in T} such that FsFt\mathcal{F}_s \subseteq \mathcal{F}_t whenever sts \leq t, with FtF\mathcal{F}_t \subseteq \mathcal{F} for all tt.[59] This increasing structure captures the non-decreasing nature of information accumulation, where events measurable at earlier times remain measurable later. Filtrations are often assumed to be right-continuous, meaning Ft=u>tFu\mathcal{F}_t = \bigcap_{u > t} \mathcal{F}_u for each tTt \in T, ensuring that the information at time tt includes all limits of information from slightly later times; this property is crucial for handling limits in stochastic models.[59] A stochastic process {Xt}tT\{X_t\}_{t \in T} defined on this filtered probability space is said to be adapted to the filtration {Ft}tT\{\mathcal{F}_t\}_{t \in T} if, for every tTt \in T, the random variable Xt:ΩSX_t: \Omega \to S (where SS is the state space) is Ft\mathcal{F}_t-measurable.[59] Adaptivity formalizes the idea that the value of the process at time tt depends only on the information available up to tt, preventing anticipation of future events. For instance, the Wiener process (standard Brownian motion) is typically defined to be adapted to its natural filtration, ensuring that its increments reveal information progressively without foreknowledge.[59] The natural filtration generated by a stochastic process {Xt}tT\{X_t\}_{t \in T} is the smallest filtration to which the process is adapted, defined as FtX=σ(Xs:st)\mathcal{F}_t^X = \sigma(X_s : s \leq t), the σ\sigma-algebra generated by all random variables XsX_s for sts \leq t.[59] This filtration encodes precisely the information revealed by the process itself up to time tt, making it fundamental for analyzing self-contained dynamics. For more refined notions of information flow, especially in preparation for stochastic integration, predictability distinguishes processes based on their measurability properties relative to the filtration. A process is progressively measurable if, for every t>0t > 0, the map (s,ω)Xs(ω)(s, \omega) \mapsto X_s(\omega) from [0,t]×Ω[0, t] \times \Omega to R\mathbb{R} is measurable with respect to the product σ\sigma-algebra B([0,t])Ft\mathcal{B}([0, t]) \otimes \mathcal{F}_t, implying adaptivity and joint measurability over finite intervals; this ensures the process can be approximated by simple functions for integration purposes.[60] Predictability, a stronger condition, requires the process to be measurable with respect to the predictable σ\sigma-algebra P\mathcal{P}, generated by left-continuous adapted processes (or equivalently, stochastic intervals [[0,τ[)[[0, \tau[) for stopping times τ\tau); optional measurability, in contrast, is with respect to the optional σ\sigma-algebra generated by right-continuous adapted processes.[60] These concepts—progressive for broad integration and predictable for avoiding jumps at unpredictable times—are essential for defining Itô integrals and handling discontinuities in paths.[60]

Modifications and Versions

In the theory of stochastic processes, two processes X=(Xt)tTX = (X_t)_{t \in T} and Y=(Yt)tTY = (Y_t)_{t \in T} defined on the same probability space are said to be modifications of each other if they possess identical finite-dimensional distributions, meaning that for any finite collection of times t1,,tnTt_1, \dots, t_n \in T and Borel sets B1,,BnB_1, \dots, B_n, the probability P(Xt1B1,,XtnBn)=P(Yt1B1,,YtnBn)P(X_{t_1} \in B_1, \dots, X_{t_n} \in B_n) = P(Y_{t_1} \in B_1, \dots, Y_{t_n} \in B_n) holds.[61] This equivalence in law allows modifications to differ in their sample paths, as the joint distributions at fixed times do not constrain the behavior between those times or the precise path realizations, provided the marginal and joint laws remain unchanged.[46] For instance, the standard Wiener process admits multiple modifications, such as one with continuous paths and another without, yet all share the same finite-dimensional Gaussian distributions with mean zero and covariance min(t,s)\min(t,s).[62] Within the class of modifications, a version of XX is a process YY such that P(Xt=Yt)=1P(X_t = Y_t) = 1 for every tTt \in T. A stronger notion is indistinguishability, where YY is indistinguishable from XX if P({ωΩ:Xt(ω)=Yt(ω) tT})=1P\left( \{\omega \in \Omega : X_t(\omega) = Y_t(\omega) \ \forall t \in T \} \right) = 1, meaning the sample paths coincide almost surely. For processes with regular paths, such as continuous or separable ones, indistinguishability is equivalent to the paths being equal almost everywhere with respect to Lebesgue measure on TT almost surely, under suitable measurability conditions.[63] To achieve uniqueness and facilitate analysis, particularly in applications involving filtrations or integrals, a regular modification is often selected by choosing a right-continuous version of the process. A right-continuous version possesses paths that are right-continuous at every time tTt \in T, with limstXs=Xt\lim_{s \downarrow t} X_s = X_t almost surely for all tt, and typically includes left limits where appropriate (càdlàg paths). This choice is possible for many classes of processes, such as Lévy processes or martingales, under conditions like those in the Kolmogorov continuity theorem, ensuring a unique representative within the equivalence class of modifications while preserving the finite-dimensional distributions.[3] Such regular versions are essential for theorems on stopping times and optional sampling, as they guarantee path regularity without altering the underlying probabilistic structure.[64]

Independence and Dependence Measures

In stochastic processes, independence is fundamentally defined in terms of σ-algebras generated by the process components. Two sub-σ-algebras F\mathcal{F} and G\mathcal{G} of the underlying probability space (Ω,F,P)(\Omega, \mathcal{F}, P) are independent if, for every AFA \in \mathcal{F} and BGB \in \mathcal{G}, P(AB)=P(A)P(B)P(A \cap B) = P(A) P(B).[65] This extends to processes: a stochastic process {Xt}\{X_t\} has independent increments if the σ-algebras generated by the increments XtkXtk1X_{t_k} - X_{t_{k-1}} over disjoint time intervals [tk1,tk][t_{k-1}, t_k] are independent.[66] For instance, the Wiener process exhibits independent increments over non-overlapping intervals.[59] Uncorrelatedness provides a weaker measure of dependence, focusing on second moments rather than full distributional properties. For components of stochastic processes, such as XtX_t and YsY_s (which may belong to the same or different processes), uncorrelatedness holds if E[(Xtμt)(Ysμs)]=0\mathbb{E}[(X_t - \mu_t)(Y_s - \mu_s)] = 0 for tst \neq s, where μt=E[Xt]\mu_t = \mathbb{E}[X_t] and μs=E[Ys]\mu_s = \mathbb{E}[Y_s].[67] In the context of a single process with zero mean, this simplifies to the increments being uncorrelated if their covariances vanish over disjoint intervals.[66] Orthogonality is a concept from the Hilbert space L2(Ω,F,P)L^2(\Omega, \mathcal{F}, P), where random variables with finite second moments form an inner product space with X,Y=E[XY]\langle X, Y \rangle = \mathbb{E}[XY]. Two such elements XX and YY (typically centered) are orthogonal if X,Y=0\langle X, Y \rangle = 0.[68] For stochastic processes, this applies to increments: a process has orthogonal increments if E[(XtXs)(XuXv)]=0\mathbb{E}[(X_t - X_s)(X_u - X_v)] = 0 whenever the intervals [s,t][s, t] and [u,v][u, v] are disjoint.[68] Independence implies uncorrelatedness (and hence orthogonality when centered) for L2L^2 random variables, as E[XY]=E[X]E[Y]\mathbb{E}[XY] = \mathbb{E}[X] \mathbb{E}[Y] under independence, yielding zero covariance.[69] The converse fails: uncorrelatedness does not imply independence. A counterexample involves ZN(0,1)Z \sim \mathcal{N}(0,1) and independent WW taking values ±1\pm 1 with equal probability 1/21/2; set X=ZX = Z and Y=WZY = W Z. Then Cov(X,Y)=E[WZ2]=E[W]E[Z2]=01=0\mathrm{Cov}(X, Y) = \mathbb{E}[W Z^2] = \mathbb{E}[W] \mathbb{E}[Z^2] = 0 \cdot 1 = 0, but XX and YY are dependent since Y=X|Y| = |X| almost surely.[69] For joint uniform distributions on [1,1]×[1,1][-1,1] \times [-1,1] restricted to the unit circle (via polar coordinates), the variables are uncorrelated but their joint distribution is singular with respect to the product measure.[69]

Regularity Conditions

Regularity conditions impose structural constraints on stochastic processes to guarantee that their sample paths exhibit desirable properties almost surely, facilitating analysis and ensuring measurability in appropriate function spaces. These conditions are essential for distinguishing processes with smooth trajectories from those with jumps or irregularities, and they often rely on the existence of suitable modifications or versions of the process. For instance, the Wiener process serves as a canonical example satisfying strong regularity, with paths that are continuous almost surely. Separability is a fundamental regularity condition that ensures a stochastic process admits a version where the path values are determined by their behavior on a countable dense subset of the index set. Specifically, for a process {Xt:tT}\{X_t : t \in T\} with TRT \subset \mathbb{R} uncountable, separability requires the existence of a countable dense set DTD \subset T such that for almost every ω\omega, the values Xt(ω)X_t(\omega) for tTt \in T are fully determined by the restriction to DD, up to a null set of paths. This property, introduced by Doob, implies that every stochastic process has a separable modification, which is crucial for avoiding pathological behaviors in uncountable index sets and ensuring the process is measurable with respect to the product σ\sigma-algebra. Continuity conditions focus on the almost sure continuity of sample paths, often quantified through bounds on the modulus of continuity. A process has continuous paths if, for almost every realization, the mapping tXt(ω)t \mapsto X_t(\omega) is continuous on TT. To establish such versions, the Kolmogorov continuity theorem provides a sufficient criterion: if there exist positive constants C,α,βC, \alpha, \beta with α>0\alpha > 0 and β>0\beta > 0 such that E[XtXsα]Ctsd+β\mathbb{E}[|X_t - X_s|^\alpha] \leq C |t - s|^{d + \beta} for all s,tTs, t \in T in a dd-dimensional setting, then the process admits a continuous modification. This theorem, originally due to Kolmogorov, enables the construction of continuous versions for processes like Brownian motion by controlling the expected increments. For processes exhibiting jumps, such as those in queueing theory or financial modeling, càdlàg (right-continuous with left limits) paths provide a weaker but still regular structure. A process has càdlàg paths almost surely if, for almost every ω\omega, the function tXt(ω)t \mapsto X_t(\omega) is right-continuous at every tTt \in T and admits finite left limits as sts \uparrow t. This property accommodates discontinuities while ensuring the paths are bounded variation or semimartingale-like in compact intervals, as formalized in the theory of stochastic integration. Càdlàg versions exist under mild conditions on the finite-dimensional distributions, making them suitable for jump-diffusion models.

Advanced Stochastic Processes

Markov Processes

A Markov process is a stochastic process that satisfies the Markov property, meaning that the conditional distribution of the future state given the entire history up to the present is determined solely by the current state. Formally, for a stochastic process (Xt)t0(X_t)_{t \geq 0} with state space EE and natural filtration (Ft)t0(\mathcal{F}_t)_{t \geq 0}, the Markov property states that for any s>0s > 0, Borel set AEA \subseteq E, and t0t \geq 0,
P(Xt+sAFt)=P(Xt+sAXt)almost surely. \mathbb{P}(X_{t+s} \in A \mid \mathcal{F}_t) = \mathbb{P}(X_{t+s} \in A \mid X_t) \quad \text{almost surely}.
This memoryless property implies that the process "forgets" its past beyond the current position, simplifying the analysis of its evolution. The transition probabilities of a Markov process encode this dependence on the current state. For a time-homogeneous Markov process starting at xEx \in E, the transition kernel is defined as Pt(x,A)=P(XtAX0=x)P_t(x, A) = \mathbb{P}(X_t \in A \mid X_0 = x) for t0t \geq 0 and Borel AEA \subseteq E. These kernels form a semigroup under composition: Ps+t=PsPtP_{s+t} = P_s P_t for all s,t0s, t \geq 0, where the product denotes the operator (PsPtf)(x)=EPs(x,dy)f(y)(P_s P_t f)(x) = \int_E P_s(x, dy) f(y) for bounded measurable functions f:ERf: E \to \mathbb{R}. This semigroup structure arises directly from the Markov property and enables the representation of the process's dynamics via functional equations.[70] A key consequence of the semigroup property is the Chapman-Kolmogorov equation, which expresses the transition probability over an interval as an integral over intermediate states:
Ps+t(x,A)=EPs(x,dy)Pt(y,A),s,t0. P_{s+t}(x, A) = \int_E P_s(x, dy) P_t(y, A), \quad s, t \geq 0.
This equation, independently derived by Chapman in 1928 and Kolmogorov in 1931, is fundamental for solving the forward and backward equations governing the evolution of transition densities in continuous-state cases. It holds for both discrete- and continuous-time Markov processes and underpins the analytical methods for their study.[71][72] Examples of Markov processes abound in probability theory. In discrete time, a Markov chain on a countable state space evolves according to fixed transition probabilities between states, as introduced by Markov in his 1906 work on sequences of dependent trials.[73] In continuous time and space, diffusion processes such as Brownian motion (Wiener process) and the Poisson process satisfy the Markov property; the former models random walks with continuous paths, while the latter counts events in fixed intervals with stationary increments. The strong Markov property extends the standard Markov property to hold at random stopping times τ\tau, which are Ft\mathcal{F}_t-adapted random variables with almost sure finite values. Specifically, for any stopping time τ\tau and s>0s > 0,
P(Xτ+sAFτ)=P(Xτ+sAXτ)almost surely on {τ<}. \mathbb{P}(X_{\tau + s} \in A \mid \mathcal{F}_\tau) = \mathbb{P}(X_{\tau + s} \in A \mid X_\tau) \quad \text{almost surely on } \{\tau < \infty\}.
This stronger version, developed by Doob in the 1950s, is crucial for processes like Brownian motion and allows restarts at unpredictable times, facilitating applications in optional sampling and decomposition theorems.

Martingales

A martingale is a stochastic process that models a sequence of random variables where the expected value of the next observation, conditional on all prior observations, equals the current value, embodying the notion of a fair game in probability theory. Formally, given a probability space (Ω,F,P)(\Omega, \mathcal{F}, P) and a filtration {Ft}tT\{\mathcal{F}_t\}_{t \in T} (where TT is a totally ordered set, often [0,)[0, \infty) or N\mathbb{N}), a stochastic process {Xt}tT\{X_t\}_{t \in T} is a martingale if it is adapted to the filtration (i.e., XtX_t is Ft\mathcal{F}_t-measurable for each tt), E[Xt]<E[|X_t|] < \infty for all tTt \in T, and satisfies the martingale property
E[XtFs]=Xsalmost surely E[X_t \mid \mathcal{F}_s] = X_s \quad \text{almost surely}
for all s<ts < t in TT. This definition was introduced by Joseph L. Doob in his foundational work on the regularity properties of families of chance variables, where martingales were first formalized as tools to study convergence and boundedness in stochastic systems. Submartingales and supermartingales extend the martingale concept to processes with directional biases in their conditional expectations. A process {Xt}\{X_t\} is a submartingale if it is adapted, integrable, and E[XtFs]XsE[X_t \mid \mathcal{F}_s] \geq X_s almost surely for s<ts < t; conversely, it is a supermartingale if E[XtFs]XsE[X_t \mid \mathcal{F}_s] \leq X_s almost surely for s<ts < t. Every martingale is both a submartingale and a supermartingale, but the inequalities allow modeling scenarios with positive or negative drifts, such as in gambling systems with house edges. These generalizations were systematically developed by Doob to analyze broader classes of stochastic processes beyond strict fairness. The Doob decomposition theorem provides a canonical way to break down submartingales into martingale and predictable components, revealing underlying structures in stochastic evolution. Specifically, for a submartingale {Xt}\{X_t\} with respect to {Ft}\{\mathcal{F}_t\}, there exists a unique decomposition Xt=Mt+AtX_t = M_t + A_t almost surely for each tt, where {Mt}\{M_t\} is a martingale, {At}\{A_t\} is a predictable process (measurable with respect to the predictable sigma-algebra generated by the filtration) that is non-decreasing and non-negative with A0=0A_0 = 0, and both processes start from the same initial value as X0X_0. This theorem, established by Doob, enables the isolation of the "noise" (martingale part) from the "trend" (predictable part), facilitating applications in decomposition and prediction. The simple symmetric random walk on the integers serves as a basic discrete-time example of a martingale, where the position after each step has conditional expectation equal to the current position. Martingales possess strong convergence properties that underpin their utility in limit theorems for stochastic processes. Doob's martingale convergence theorem states that if {Xn}nN\{X_n\}_{n \in \mathbb{N}} is a martingale (or more generally, a submartingale) satisfying supnE[Xn]<\sup_n E[|X_n|] < \infty, then XnX_n converges almost surely to a random variable XL1X_\infty \in L^1 as nn \to \infty, with E[X]supnE[Xn]E[|X_\infty|] \leq \sup_n E[|X_n|]. This result was originally proved by Doob for discrete-time cases using upcrossing inequalities to control oscillations. For L^1-convergence, uniform integrability of {Xn}\{X_n\}—meaning supnE[Xn1{Xn>K}]0\sup_n E[|X_n| \mathbf{1}_{\{|X_n| > K\}}] \to 0 as KK \to \infty—is required, ensuring E[XnX]0E[|X_n - X_\infty|] \to 0. Extensions to continuous time follow under right-continuity assumptions on the paths, preserving the almost sure convergence to an integrable limit.

Lévy Processes

A Lévy process is a stochastic process (Xt)t0(X_t)_{t \geq 0} with values in Rd\mathbb{R}^d, starting at X0=0X_0 = 0 almost surely, that possesses stationary and independent increments, right-continuous paths with left limits (càdlàg paths), and stochastic continuity, meaning limt0P(XtX0>ϵ)=0\lim_{t \to 0} P(|X_t - X_0| > \epsilon) = 0 for every ϵ>0\epsilon > 0.[74] The stationary increments property implies that the distribution of Xs+tXsX_{s+t} - X_s depends only on tt, while independence ensures that increments over disjoint intervals are independent random variables.[74] This structure generalizes classical processes like the Wiener process and Poisson process, which satisfy these conditions as special cases.[74] The characteristic function of a Lévy process provides a complete description of its law through the Lévy–Khintchine formula. For XtX_t, it is given by
E[eiuXt]=exp(tψ(u)), \mathbb{E}[e^{i u \cdot X_t}] = \exp\left(t \psi(u)\right),
where uRdu \in \mathbb{R}^d and the characteristic exponent ψ(u)\psi(u) takes the form
ψ(u)=ibu12uΣu+Rd{0}(eiux1iux1x<1)ν(dx). \psi(u) = i b \cdot u - \frac{1}{2} u^\top \Sigma u + \int_{\mathbb{R}^d \setminus \{0\}} \left( e^{i u \cdot x} - 1 - i u \cdot x \mathbf{1}_{|x| < 1} \right) \nu(dx).
Here, bRdb \in \mathbb{R}^d is the drift vector, Σ\Sigma is a symmetric positive semidefinite diffusion matrix capturing the continuous Gaussian component, and ν\nu is the Lévy measure describing the jumps, satisfying Rd{0}(1x2)ν(dx)<\int_{\mathbb{R}^d \setminus \{0\}} (1 \wedge |x|^2) \nu(dx) < \infty.[75] This triplet (b,Σ,ν)(b, \Sigma, \nu) uniquely determines the process among Lévy processes with the same filtration.[75] Prominent examples of Lévy processes include Brownian motion with drift, where ν=0\nu = 0 and Σ\Sigma is positive definite, yielding continuous paths; the compound Poisson process, characterized by a finite Lévy measure ν\nu concentrated on jumps of finite activity; and stable Lévy processes, which have self-similar increments with heavy tails when Σ=0\Sigma = 0 and ν\nu follows a stable form.[74] These examples illustrate the broad class, encompassing both continuous and jump components.[74] The increments of a Lévy process are infinitely divisible, meaning for each t>0t > 0, the distribution of XtX_t can be expressed as the convolution of nn identical distributions for any nNn \in \mathbb{N}.[74] Conversely, every infinitely divisible distribution arises as the law of X1X_1 for some Lévy process.[74] This property allows representation of general Lévy increments as limits of compound Poisson processes, where the jump measure ν\nu is truncated and approximated by finite-activity jumps, converging in distribution as the truncation refines.[74]

Point Processes and Random Fields

Point processes represent a class of stochastic processes that model random configurations of points in a general measurable space, often viewed as random counting measures NN on that space. Unlike standard processes indexed by time, point processes capture discrete events or locations without inherent order, generalizing concepts like the one-dimensional Poisson process to higher-dimensional or abstract settings. A prominent example is the Poisson point process, defined on a space SS with intensity measure Λ\Lambda, where the number of points in any bounded region follows a Poisson distribution with mean Λ\Lambda of that region, and counts in disjoint regions are independent. A key result for such processes is Campbell's theorem, which states that for a non-negative measurable function ff,
E[xNf(x)]=Sf(x)Λ(dx), \mathbb{E}\left[ \sum_{x \in N} f(x) \right] = \int_S f(x) \, \Lambda(dx),
providing the expected value of sums over the points via the intensity measure. This theorem facilitates moment calculations and is foundational for analyzing functionals of point processes. Palm distributions offer a conditional perspective on point processes, particularly for stationary cases, by describing the distribution of the process given the presence of a point at a specific location, such as the origin.[76] Formally, the reduced Palm distribution conditions on points at designated locations while removing those points from the configuration, enabling the study of typical structures around observed events; this concept originated in Conrad Palm's 1943 analysis of telephone traffic fluctuations.[76] Random fields extend stochastic processes to multi-dimensional index sets TT, such as spatial domains in Rd\mathbb{R}^d, where the process X:T×ΩEX: T \times \Omega \to E assigns random values to each point in TT.[77] These fields are crucial for modeling phenomena with spatial dependence, often assuming isotropy, where statistical properties like the covariance function depend only on the distance between points, C(ri,rj)=C(rirj)C(\mathbf{r}_i, \mathbf{r}_j) = C(|\mathbf{r}_i - \mathbf{r}_j|).[78] Gaussian random fields, a widely studied class, have finite-dimensional distributions that are multivariate normal, fully specified by mean and covariance functions, and exhibit properties like continuity and smoothness under suitable conditions on the covariance. They are prevalent in spatial statistics for interpolating unobserved values via kriging. Gibbs random fields, on the other hand, are defined through Gibbs measures that satisfy the Dobrushin-Lanford-Ruelle equations, incorporating local interaction potentials to model dependent lattice or continuous configurations in statistical mechanics and spatial analysis.

Mathematical Construction

Challenges in Defining Processes

Defining a stochastic process on continuous index sets, such as the real line, presents significant challenges due to the infinite-dimensional nature of the path space. While finite-dimensional distributions (f.d.d.) provide a natural starting point for specification, extending these to a consistent probability measure on the full path space requires careful conditions to avoid inconsistencies or pathological behaviors. In general measurable spaces, consistent f.d.d. do not always admit an extension to a probability measure on the product sigma-algebra, as demonstrated by counterexamples where the cylinder sets fail to generate a well-defined process. A key issue arises in the measurability of sample paths. Without additional regularity assumptions, such as right-continuity or bounded variation, the paths of a stochastic process defined via f.d.d. may not be measurable functions from the probability space to the path space equipped with the Borel sigma-algebra. This non-measurability complicates the analysis of path properties and integrals, necessitating the imposition of conditions like cadlag (right-continuous with left limits) to ensure almost sure measurability. The problem stems from the fact that the natural sigma-algebra on the path space, generated by cylinders, may not capture the full Borel structure for uncountable index sets, leading to potential gaps in the probabilistic framework. Further difficulties emerge when considering convergence of processes or tightness of measure families. For the path space to support useful weak convergence results, it must typically be a Polish space—a complete separable metric space—to leverage Prohorov's theorem, which equates tightness of probability measures with relative compactness in the weak topology. In non-Polish settings, such as arbitrary product spaces over continuous time, tightness may fail to imply compactness, hindering the construction of limiting processes and requiring auxiliary topologies like Skorokhod for resolution. This topological requirement underscores the need for complete separable metric structures to guarantee the existence and well-behaved properties of stochastic processes on continuous domains. Historically, these definitional hurdles were illuminated by paradoxes revealing the limitations of naive extensions. For instance, early attempts to define processes with continuous paths encountered issues where consistent f.d.d. could not be realized by measurable paths without invoking specific metric assumptions, prompting the development of regularity conditions derived from key probabilistic properties like continuity in probability. Such insights have shaped the rigorous foundations of stochastic processes, emphasizing the interplay between measure-theoretic consistency and topological completeness.

Canonical Spaces and Measure Constructions

In the construction of stochastic processes, the canonical space serves as the natural sample space for realizing the process paths. For a stochastic process (Xt)tT(X_t)_{t \in T} with state space SS and time index set TT, the canonical path space is the set STS^T of all functions from TT to SS, often equipped with the product topology. The σ-algebra on this space is the cylinder σ-algebra, generated by the finite-dimensional cylinders {ωST:(Xt1(ω),,Xtn(ω))B}\{ \omega \in S^T : (X_{t_1}(\omega), \dots, X_{t_n}(\omega)) \in B \} for finite subsets {t1,,tn}T\{t_1, \dots, t_n\} \subset T and Borel sets BSnB \subset S^n. This structure ensures that the finite-dimensional distributions (f.d.d.s) determine the measurable properties of the process.[18] A prominent example of a canonical space is the Wiener space for Brownian motion, defined as C[0,)C[0, \infty), the space of continuous functions ω:[0,)R\omega: [0, \infty) \to \mathbb{R} with ω(0)=0\omega(0) = 0, under the supremum norm on compact intervals. The Wiener measure W\mathbb{W} is the unique probability measure on the Borel σ-algebra of this space such that the coordinate process Wt(ω)=ω(t)W_t(\omega) = \omega(t) is a standard Brownian motion, satisfying the properties of continuous paths, independent Gaussian increments with mean zero and variance tt, starting at zero. This measure is constructed to resolve the challenges of defining processes with specified f.d.d.s on infinite-dimensional spaces.[3] The Kolmogorov extension theorem provides the foundational tool for constructing probability measures on these canonical spaces. Given a consistent family of probability measures {μn}nN\{\mu_n\}_{n \in \mathbb{N}} on the finite products SnS^n, where consistency means that for any m<nm < n and indices i1,,im{1,,n}i_1, \dots, i_m \in \{1, \dots, n\}, the marginal μn\mu_n on the i1,,imi_1, \dots, i_m-coordinates equals μm\mu_m, there exists a unique probability measure μ\mu on the product σ-algebra of STS^T such that the f.d.d.s of μ\mu match the μn\mu_n. This theorem guarantees the existence of a stochastic process with prescribed consistent f.d.d.s, bridging finite-dimensional specifications to the full path measure.[79] To ensure the existence of processes with desirable convergence properties, such as weak convergence of measures on path spaces, tightness plays a crucial role. The Prokhorov criterion characterizes tightness: a family of probability measures {Pα}\{\mathbb{P}_\alpha\} on a metric space is tight if, for every ϵ>0\epsilon > 0, there exists a compact set KK such that Pα(K)1ϵ\mathbb{P}_\alpha(K) \geq 1 - \epsilon for all α\alpha. On complete separable metric spaces (Polish spaces), tightness implies that every sequence in the family has a weakly convergent subsequence, with the limit measure also in the closure of the family. This criterion is essential for verifying the relative compactness of sequences of process measures in applications involving weak convergence.[80] For specific classes like Lévy processes, existence follows from the structure of their characteristic functions. A Lévy process has stationary independent increments with almost surely right-continuous paths with left limits, and its one-dimensional distributions are infinitely divisible. The Lévy-Khintchine formula represents the characteristic function ϕt(u)=E[eiuXt]=exp{tψ(u)}\phi_t(u) = \mathbb{E}[e^{i u X_t}] = \exp\{t \psi(u)\}, where ψ(u)=ibu12σ2u2+R{0}(eiux1iux1x<1)ν(dx)\psi(u) = i b u - \frac{1}{2} \sigma^2 u^2 + \int_{\mathbb{R} \setminus \{0\}} (e^{i u x} - 1 - i u x \mathbf{1}_{|x|<1}) \nu(dx) for drift bRb \in \mathbb{R}, diffusion coefficient σ0\sigma \geq 0, and Lévy measure ν\nu. This form ensures consistency of the f.d.d.s via the independent increments property, allowing application of the Kolmogorov extension to construct the process measure on the canonical space D[0,)\mathbb{D}[0, \infty) of càdlàg functions.[81]

Skorokhod Topology and Convergence

The Skorokhod space, denoted D[0,)D[0,\infty), consists of all real-valued functions on [0,)[0,\infty) that are right-continuous with left limits (càdlàg) everywhere, providing a natural setting for modeling stochastic processes with possible jumps, such as those arising in queueing theory or financial modeling. This space is equipped with the Skorokhod topology, which is generated by a metric that accounts for both the spatial distance between functions and a time reparameterization to handle discontinuities. Specifically, the metric d(X,Y)d(X,Y) between two functions X,YD[0,)X, Y \in D[0,\infty) is defined as the infimum over all continuous, strictly increasing time-change functions λ:[0,)[0,)\lambda: [0,\infty) \to [0,\infty) with λ(0)=0\lambda(0)=0 of XYλ+λid\|X - Y \circ \lambda\| + \|\lambda - \mathrm{id}\|, where \|\cdot\| denotes the supremum norm adjusted for finite intervals (often via supT>0min(1,dT(X,Y))\sup_{T>0} \min(1, d_T(X,Y)) for compactness on [0,T][0,T]). This construction, introduced by A.V. Skorokhod, ensures the space is complete and separable, making it suitable for probabilistic limits despite the lack of uniform continuity in paths. Convergence in the Skorokhod topology is particularly useful for weak convergence of probability measures on D[0,)D[0,\infty), known as convergence in distribution for stochastic processes. A sequence of processes XnX_n converges in distribution to XX if the measures PXn\mathbb{P}_{X_n} converge weakly to PX\mathbb{P}_X in this topology, which requires tightness of {PXn}\{\mathbb{P}_{X_n}\} and convergence of finite-dimensional distributions at continuity points of the limit. Unlike the uniform topology on continuous functions, the Skorokhod metric permits small time distortions, allowing convergence even when jump times in XnX_n do not align exactly with those in XX, provided the jumps are of finite activity. This weak convergence framework is essential for establishing functional limit theorems, as it preserves probabilistic structure under scaling. A key application is in functional limit theorems, such as invariance principles that approximate discrete processes by continuous limits. For instance, Donsker's invariance principle states that the scaled random walk Sn(t)=n1/2k=1ntξkS_n(t) = n^{-1/2} \sum_{k=1}^{\lfloor nt \rfloor} \xi_k, where ξk\xi_k are i.i.d. with mean zero and finite variance, converges in distribution in the Skorokhod topology on D[0,1]D[0,1] to a standard Brownian motion W(t)W(t). This result extends to D[0,)D[0,\infty) by considering restrictions to compact intervals, highlighting how the topology bridges discrete and continuous path behaviors. The principle relies on the Skorokhod metric's flexibility, as the polygonal paths of the random walk converge to the continuous Brownian paths despite minor time-warping near jumps (which are absent in the limit). The distinction between path continuity and the Skorokhod metric underscores its utility: while càdlàg paths in D[0,)D[0,\infty) may have discontinuities, the topology induces uniform convergence on compact sets when the limit process has continuous paths, as continuous functions are dense in the space. If XnXX_n \to X in Skorokhod topology and XX is continuous, then the convergence is actually uniform in probability, i.e., suptXn(t)X(t)0\sup_t |X_n(t) - X(t)| \to 0 in probability. Conversely, for discontinuous limits like Lévy processes, the metric's time-reparameterization is crucial to capture asymptotic behavior without requiring exact synchronization of jumps. This balance makes the Skorokhod topology indispensable for modern stochastic analysis, enabling rigorous limits in non-smooth settings.

Historical Development

Origins in Probability and Statistics

The foundations of stochastic processes emerged from early probability theory in the 17th century, driven by efforts to analyze games of chance and repeated random events. Christiaan Huygens's 1657 treatise De Ratiociniis in Ludo Aleae marked the first systematic application of mathematics to gambling problems, introducing the concept of expected value as a fair price for random outcomes and establishing rules for dividing stakes in interrupted games, which implicitly modeled sequences of probabilistic trials.[82] This work built on the 1654 correspondence between Blaise Pascal and Pierre de Fermat, who resolved the "problem of points" by deriving probabilities for incomplete games through combinatorial enumeration, laying groundwork for handling dependent sequential events.[83] Jacob Bernoulli advanced these ideas in his posthumously published Ars Conjectandi (1713), which formalized the analysis of repeated independent trials—now known as the Bernoulli process—and proved the law of large numbers, demonstrating that the average of outcomes from many trials converges to the expected value with high probability.[84] Bernoulli's theorem provided a rigorous basis for viewing sequences of random events as predictable in the aggregate, influencing later conceptions of stochastic sequences. In the 19th century, Siméon Denis Poisson extended probabilistic modeling to legal and social contexts in Recherches sur la probabilité des jugements en matière criminelle et en matière civile (1837), where he derived the Poisson distribution as a limit law for rare events in large numbers of independent trials, capturing the probability of event counts over time intervals.[85] This distribution became essential for describing processes with sporadic occurrences, bridging discrete trials to continuous-time randomness. The late 19th century saw probability intertwined with statistical mechanics, as physicists sought to explain macroscopic phenomena through microscopic random motions. Ludwig Boltzmann's papers in the 1870s, including his derivation of the Maxwell-Boltzmann distribution, employed probabilistic ensembles to model gas particle collisions and velocities, showing how irreversible thermodynamic laws arise from reversible microscopic dynamics averaged over random states.[86] J. Willard Gibbs synthesized these approaches in Elementary Principles in Statistical Mechanics (1902), introducing the Gibbs ensemble and phase space probability densities to predict system evolution under random fluctuations, formalizing the statistical foundation for dynamic processes.[87] Early 20th-century developments included Louis Bachelier's 1900 doctoral thesis, which modeled stock price fluctuations as a random walk (Brownian motion) for financial applications, and Albert Einstein's 1905 explanation of physical Brownian motion as diffusion due to molecular collisions, providing a mathematical framework for continuous stochastic paths.[88][89] The Wiener process later drew physical roots from such Brownian motion descriptions in gases. Specific models of random displacement soon followed. Karl Pearson posed the "random walk" problem in 1905, modeling the net displacement after a series of equal random steps in one or two dimensions to approximate diffusive paths, with solutions revealing Gaussian limiting distributions for large steps. In 1907, Paul and Tatyana Ehrenfest introduced the "dog-flea" model—two dogs exchanging fleas randomly—to illustrate molecular diffusion and approach to equilibrium, demonstrating how stochastic transfers between compartments lead to binomial equilibrium distributions.[90] These early constructs highlighted the utility of random processes in capturing irregular yet statistically regular behaviors.

Contributions from Measure Theory

The axiomatic foundation of probability theory, established through measure-theoretic principles in the early 1930s, provided the rigorous framework necessary for defining stochastic processes as measurable functions on probability spaces. Andrei Kolmogorov's seminal 1933 monograph, Grundbegriffe der Wahrscheinlichkeitsrechnung, introduced probability as a special case of measure theory, where events correspond to measurable sets and probabilities to measures on a sigma-algebra, enabling the treatment of infinite sequences of random variables central to stochastic processes.[91] This measure-theoretic approach resolved earlier heuristic ambiguities in process definitions by ensuring consistency and measurability, allowing stochastic processes to be viewed as coordinate mappings from abstract spaces to time-indexed outcomes.[92] Building on this foundation, the 1930s saw the development of extension theorems that guaranteed the existence of stochastic processes from consistent families of finite-dimensional distributions. Kolmogorov's extension theorem, articulated in his 1933 work and subsequent elaborations, demonstrated that a collection of probability measures on finite-dimensional Euclidean spaces, satisfying consistency conditions (such as marginal agreement), could be uniquely extended to a measure on the space of all sample paths, thus constructing the process on a canonical probability space.[93] This theorem addressed key challenges in defining processes over uncountable index sets, like continuous time, by leveraging Kolmogorov's axioms to ensure the extended measure is sigma-additive and complete.[94] In the 1940s, Joseph L. Doob advanced the measure-theoretic treatment of stochastic processes through his development of martingale theory and its connections to potential theory. Doob's work, beginning with papers in the early 1940s, reformulated martingales as processes satisfying the conditional expectation property with respect to filtrations defined via measures, providing tools for convergence and decomposition results in general spaces.[95] His integration of these concepts into potential theory used harmonic functions adapted to measure spaces, enabling the analysis of sub- and super-martingales as solutions to boundary value problems in probabilistic terms.[96] Paul Lévy's contributions in the 1940s further solidified the measure-theoretic underpinnings of stochastic processes, particularly through advancements in stochastic integration and path decompositions. In works such as his 1948 monograph Processus Stochastiques et Mouvement Brownien, Lévy extended integration techniques to non-differentiable paths using measure-theoretic limits and occupation times, allowing for the rigorous handling of irregular sample functions. His decompositions, including those separating continuous and jump components in processes with independent increments, relied on characteristic functions and Lévy measures to classify path behaviors within abstract probability spaces.[97]

Mid-20th Century Advances and Key Figures

In the post-World War II era, stochastic processes advanced significantly through applications in signal processing and foundational theoretical frameworks. Norbert Wiener's development of the Wiener filter in the 1940s provided a cornerstone for optimal estimation in noisy environments, particularly for predicting stationary time series in engineering contexts such as anti-aircraft control systems. This work, formalized in his 1949 monograph, introduced linear prediction methods based on spectral analysis of stochastic signals, influencing subsequent developments in time-series analysis.[98] Joseph L. Doob's 1953 treatise Stochastic Processes systematized the field by rigorously defining processes via measure-theoretic probability, emphasizing martingales and their role in unifying discrete and continuous models. Doob's contributions, including the martingale convergence theorem, established probabilistic tools for handling randomness over time, bridging earlier work on Markov processes with modern analysis. Meanwhile, William Feller's two-volume An Introduction to Probability Theory and Its Applications (Volume I, 1950) detailed Markov chains, highlighting their irreducible and recurrent properties, and applied them to genetics, such as modeling allele frequencies under mutation and selection. Feller's exposition made these chains accessible, demonstrating their utility in simulating evolutionary dynamics. The 1960s and 1970s saw the popularization of Itô calculus, originally introduced by Kiyosi Itô in his 1944 paper on stochastic integrals with respect to Brownian motion, which enabled the differentiation of processes under quadratic variation. Itô's framework, extended through seminars and collaborations, facilitated the solution of stochastic differential equations modeling diffusion phenomena. Daniel W. Stroock and S. R. S. Varadhan's martingale problem approach, introduced in their 1969 paper, characterized diffusion processes via generator operators without requiring explicit path constructions, providing a probabilistic alternative to PDE methods influenced by measure theory contributions.[99][100] Key figures shaped these advances: Itô's stochastic calculus remains foundational for irregular paths; Henry P. McKean advanced integral representations and diffusion theory in his 1969 monograph Stochastic Integrals, co-developing tools for non-linear interactions like McKean-Vlasov equations. Daniel Revuz and Marc Yor's 1991 text Continuous Martingales and Brownian Motion synthesized martingale theory with excursions and local times, serving as a comprehensive reference for pathwise properties.[101]

Applications Across Disciplines

Finance and Risk Modeling

Stochastic processes play a central role in financial modeling by capturing the random evolution of asset prices and enabling the valuation of derivatives under uncertainty. In finance, diffusions such as Brownian motion serve as foundational building blocks for describing continuous price fluctuations, while more advanced processes incorporate volatility clustering and jumps to better reflect market dynamics. Risk-neutral pricing frameworks rely on martingales to ensure no-arbitrage conditions, allowing the adjustment of drift terms to match observed market prices.[102] A cornerstone model is geometric Brownian motion (GBM), which assumes that asset prices follow a lognormal distribution to ensure positivity. The dynamics are governed by the stochastic differential equation
dSt=μStdt+σStdWt, dS_t = \mu S_t \, dt + \sigma S_t \, dW_t,
where $ S_t $ is the asset price at time $ t $, $ \mu $ is the drift, $ \sigma $ is the volatility, and $ W_t $ is a standard Wiener process. The explicit solution is
St=S0exp((μσ22)t+σWt), S_t = S_0 \exp\left( \left( \mu - \frac{\sigma^2}{2} \right) t + \sigma W_t \right),
demonstrating exponential growth with random perturbations. This model, introduced by Samuelson for warrant pricing, posits that logarithmic returns are normally distributed, facilitating tractable simulations and analytical solutions for basic derivatives.[103][104] The Black-Scholes framework revolutionized option pricing by deriving a partial differential equation (PDE) from Itô's lemma applied to GBM under risk-neutral measure, where the drift equals the risk-free rate $ r $. The resulting closed-form formula for a European call option is
C=SN(d1)KerTN(d2), C = S N(d_1) - K e^{-rT} N(d_2),
with $ d_1 = \frac{\ln(S/K) + (r + \sigma^2/2)T}{\sigma \sqrt{T}} $ and $ d_2 = d_1 - \sigma \sqrt{T} $, where $ N(\cdot) $ is the cumulative standard normal distribution, $ K $ is the strike, and $ T $ is maturity. This approach, detailed in the seminal 1973 paper, assumes constant volatility and enables hedging strategies via dynamic replication. However, empirical evidence of volatility smiles and varying implied volatilities led to extensions incorporating stochastic volatility.[102] The Heston model addresses these limitations by allowing volatility to follow a mean-reverting square-root process, specifically the Cox-Ingersoll-Ross (CIR) diffusion for the variance $ v_t $:
dvt=κ(θvt)dt+ξvtdWtv, dv_t = \kappa (\theta - v_t) \, dt + \xi \sqrt{v_t} \, dW_t^v,
coupled with the asset dynamics $ dS_t = r S_t , dt + \sqrt{v_t} S_t , dW_t^S $, where correlation between the Brownian motions $ W^S $ and $ W^v $ captures the leverage effect. The CIR process ensures non-negative variance under Feller conditions ($ 2\kappa\theta > \xi^2 $) and was originally proposed for interest rates but adapted here for equity volatility. Heston's 1993 model yields semi-closed-form prices via Fourier inversion, improving fits to observed option surfaces during volatile periods.[105] In risk modeling, stochastic processes underpin measures like Value at Risk (VaR), which quantifies potential losses at a confidence level, often computed via Monte Carlo simulations of paths from models like GBM or Heston. Simulations generate thousands of scenarios to estimate the quantile of the portfolio loss distribution, accounting for path-dependent features in complex instruments. For instance, under GBM, returns are simulated iteratively, and VaR is the negative percentile of terminal values. This method, evaluated empirically against historical data, provides flexibility for non-normal distributions but requires computational efficiency for real-time applications. Market crashes and fat tails necessitate models with jumps, where Lévy processes generalize diffusions by adding discontinuous increments, such as compound Poisson jumps. Merton's 1976 jump-diffusion model extends GBM with Poisson-driven jumps log-normally distributed, capturing sudden price drops as seen in 1987 or 2008. The asset dynamics become $ dS_t / S_{t-} = \mu , dt + \sigma , dW_t + dJ_t $, where $ J_t $ is the jump component, allowing VaR simulations to incorporate tail risks beyond Gaussian assumptions and improving crash predictions.[106]

Physics and Engineering Systems

Stochastic processes play a central role in modeling physical phenomena involving randomness, such as particle diffusion and signal propagation in engineering systems. In physics, Brownian motion exemplifies this, describing the irregular movement of microscopic particles suspended in a fluid due to collisions with surrounding molecules. Albert Einstein provided the first quantitative theory of Brownian motion in 1905, deriving the mean squared displacement of a particle as proportional to time, which supported the atomic hypothesis of matter.[37] This model laid the foundation for understanding diffusion processes, where the particle's position follows a Gaussian distribution with variance scaling linearly with time. To capture the dynamics more explicitly, Paul Langevin introduced a stochastic differential equation in 1908 that incorporates both deterministic friction and random fluctuations. The Langevin equation is given by
dXt=γXtdt+2DdWt, dX_t = -\gamma X_t \, dt + \sqrt{2D} \, dW_t,
where XtX_t is the particle position at time tt, γ\gamma is the friction coefficient, DD is the diffusion constant, and WtW_t is a Wiener process representing the random forcing.[107] This equation models the balance between viscous drag and thermal noise, enabling simulations of particle trajectories in fluids and gases, with applications in colloid science and polymer dynamics. The Wiener process, formalized mathematically by Norbert Wiener in the 1920s, underpins these models by providing a continuous-time limit of random walks, essential for describing thermal fluctuations in physical systems.[108] In engineering, stochastic processes are vital for analyzing queueing systems, which arise in communication networks, manufacturing lines, and service operations. The M/M/1 queue models a single-server system with Poisson arrivals and exponential service times, analyzed as a continuous-time birth-death Markov chain where births represent arrivals at rate λ\lambda and deaths represent service completions at rate μ\mu.[109] The steady-state probability of nn customers in the system is πn=(1ρ)ρn\pi_n = (1 - \rho) \rho^n for utilization ρ=λ/μ<1\rho = \lambda / \mu < 1, allowing computation of metrics like average queue length. A key relation, Little's law, states that the long-run average number of customers LL equals the arrival rate λ\lambda times the average time in system WW, or L=λWL = \lambda W, proven rigorously in 1961 and applicable to stable queueing networks under mild conditions. Signal processing and control systems leverage stochastic processes for estimation in noisy environments. The Kalman filter, developed by Rudolf E. Kalman in 1960, provides an optimal recursive algorithm for estimating the state of a linear dynamic system from noisy measurements, assuming Gaussian noise modeled by stochastic processes.[110] It minimizes the mean squared error through prediction and update steps, with the state evolution following xk=Axk1+wk1x_{k} = A x_{k-1} + w_{k-1} and observations zk=Hxk+vkz_k = H x_k + v_k, where ww and vv are process and measurement noises. This has been extended to nonlinear cases via the extended Kalman filter, finding widespread use in aerospace guidance, robotics, and sensor fusion. Reliability engineering employs stochastic processes to model component failures and system availability. Failure times are often modeled as a Poisson process, where events occur at constant rate λ\lambda, implying exponentially distributed inter-failure times with memoryless property suitable for repairable systems under steady-state assumptions.[111] Renewal theory generalizes this by considering arbitrary inter-renewal distributions, tracking the number of failures over time and the age or residual life of components; for example, the renewal function m(t)m(t) gives the expected number of renewals by time tt, asymptotically m(t)t/μm(t) \sim t / \mu for mean inter-renewal μ\mu.[112] Point processes extend these ideas to model irregular event occurrences, such as defect detections in materials or seismic activities in structural engineering.

Biology and Population Modeling

Stochastic processes play a crucial role in modeling biological systems where randomness arises from demographic fluctuations, environmental variability, and individual-level events, particularly in population dynamics, ecology, genetics, and epidemiology. In biology, these models capture the inherent uncertainty in birth, death, mutation, and interaction rates, enabling predictions of extinction risks, outbreak thresholds, and evolutionary trajectories that deterministic models overlook. By incorporating stochasticity, researchers can assess the probability of rare events like population collapse or rapid disease spread, which are critical for conservation and public health strategies. Birth-death processes, as continuous-time Markov chains, model population size changes through random birth and death events, providing a foundational framework for ecological and genetic applications. In population biology, these processes describe how species abundances evolve under stochastic influences, with transition rates depending on current population size to reflect density-dependent effects. Seminal work by Kendall established the analytical foundations for computing transition probabilities and extinction probabilities in such models, highlighting their utility in forecasting long-term population viability. In genetics, the Moran model extends this to finite populations, simulating allele frequency changes via overlapping generations where individuals reproduce and die at constant rates, preserving population size while allowing genetic drift to drive fixation or loss of variants. This model has been instrumental in understanding neutral evolution and the time to fixation in small populations. The stochastic logistic model addresses density-dependent growth by incorporating environmental noise into the classic logistic equation, yielding the stochastic differential equation $ dN = r N (1 - N/K) , dt + \sigma N , dW $, where $ N $ is population size, $ r $ is the intrinsic growth rate, $ K $ is carrying capacity, $ \sigma $ quantifies noise intensity, and $ dW $ is Wiener process increment. This formulation arises from diffusion approximations of discrete birth-death processes with logistic regulation, capturing how random fluctuations can push populations toward extinction even when the deterministic mean growth is positive. Extinction risks are elevated near the Allee threshold or under high noise, with analytical approximations showing that the quasi-stationary distribution has a variance scaling with $ \sigma^2 / r $, informing conservation efforts for endangered species facing habitat stochasticity. In epidemiology, stochastic variants of the SIR (susceptible-infected-recovered) model treat transitions between compartments as Poisson-distributed events, allowing for variability in contact rates and recovery times that deterministic versions ignore. These models reveal the role of demographic stochasticity in small populations, where outbreaks may fail to ignite due to chance, with the basic reproduction number $ R_0 $ determining the supercritical branching regime for sustained transmission. Branching processes approximate early epidemic phases, modeling each infected individual as the progenitor of a random offspring distribution of secondary cases, with extinction probability solving $ s = f(s) $ where $ f $ is the probability generating function; this framework, applied to outbreaks like measles, quantifies invasion probabilities and herd immunity thresholds. Phylodynamics integrates stochastic processes to reconstruct evolutionary histories from genetic data, using coalescent processes to trace lineages backward in time through a population. Kingman's coalescent models the genealogy of a sample as a Markov process where pairs of lineages merge at rates inversely proportional to ancestral population size, assuming constant size and no selection for neutral evolution. In phylodynamics, birth-death models link forward-time population dynamics to this backward-time coalescent, enabling inference of transmission rates and sampling intensities from pathogen phylogenies, as in HIV or influenza studies where stochastic sampling through time reveals epidemic trajectories. This duality allows estimation of parameters like the effective reproduction number from tree shapes, advancing real-time surveillance of emerging diseases.

References

User Avatar
No comments yet.