[{"content":"The Machinery of Generation TL;DR\nThis 3-part introduction to flow matching is largely inspired by reference [1].\nIf you do not have time to complete the full course and companion code, this series is a fast, practical alternative: it distills the core ideas with high-quality visualizations and cleaner code to make the material easier to absorb.\n1. Generation as Sampling and Transformation The starting point of generative modeling is the sampling problem. We are given training samples\n$$ z_1,\\ldots,z_N \\sim p_{\\text{data}}, $$where each object is represented as a vector $z \\in \\mathbb{R}^d$, and the goal is to generate new samples from the same data distribution $p_{\\text{data}}$. A useful way to think about this task is the following: generation is not \u0026ldquo;creating something from nothing,\u0026rdquo; but sampling through transformation.\nMore precisely, we first sample from a simple initial distribution\n$$ p_{\\text{init}} = \\mathcal{N}(0, I_d), $$and then transform those easy samples into samples from $p_{\\text{data}}$. In this picture, the two key words are:\nSampling: begin with a point drawn from a simple distribution that we know how to sample from efficiently. Transformation: apply a rule that gradually reshapes that simple distribution into the data distribution. This perspective is the basic machinery of generation. Instead of asking a model to directly output a data sample in one step, we ask it to learn how a cloud of simple random points should move so that, by the end of the process, the cloud matches the geometry of the data distribution. Generation is therefore a transport problem: start from noise, then continuously transform that noise into data.\nA Toy Experiment To make these ideas concrete, consider a two-dimensional toy experiment that we will keep using throughout the article. In this example, the source distribution is\n$$ p_{\\text{source}} = p_{\\text{init}} = p_{\\text{simple}} = \\mathcal{N}(0, I_2), $$and the target distribution $p_{\\text{data}}$ is a symmetric two-dimensional Gaussian mixture with five modes. Because everything lives in $\\mathbb{R}^2$, the full generative process can be visualized directly.\nFigure 1. A running toy experiment: source and target distributions, learned trajectories, and the learned vector field. The accompanying code is available at [2]. The left panel shows the basic sampling problem. The red density is the source distribution $p_{\\text{source}}$, a simple Gaussian that is easy to sample from. The blue density is the target distribution $p_{\\text{data}}$, a five-mode mixture representing the data distribution we want to generate from. This panel captures the meaning of sampling and transformation: we start from easy samples, but the goal is to transform them so that they end up distributed like the target data.\nThe middle panel shows trajectories. Each black curve is the path of one particle that starts from a sampled initial point and evolves over time under the learned ODE. A trajectory therefore records how one sample moves from the source region toward one of the target modes.\nThe right panel shows the learned vector field $u_t^\\theta(x)$. Each arrow gives the instantaneous direction and speed that the model assigns to a point in space, and the arrow color indicates the magnitude $\\|u_t^\\theta(x)\\|$. The vector field is the local motion rule, while the family of all trajectories generated by this rule is the flow. Equivalently, the flow is the global transformation that carries the whole cloud of source samples toward the target distribution.\nThis toy experiment will serve as the running example for the rest of the discussion. It gives a concrete picture of generation: first sample from $p_{\\text{source}}$, then transform those samples by following the learned flow until they match $p_{\\text{data}}$.\n2. From ODEs to Flow Matching To make that transformation precise, it is formulated as an ordinary differential equation (ODE). The idea is that the desired transformation is simulated by a time-dependent dynamical system:\n$$ \\frac{d}{dt}X_t = u_t(X_t), \\qquad X_0 = x_0. $$Here $u_t(x)$ is a time-dependent velocity field that tells a particle located at $x$ how it should move at time $t$. This ODE is the mathematical object used to simulate the transformation discussed above. In the language of flow matching, this ODE defines a flow.\nThree concepts are fundamental.\nFirst, a trajectory is the solution curve obtained from one initial condition $x_0$. Once we choose the starting point, the ODE traces a path\n$$ t \\mapsto X_t $$showing how that one particle moves over time.\nSecond, the vector field $u_t(x)$ is the rule that assigns a velocity vector to every location $x$ at every time $t$. It is the local motion law of the system.\nThird, the flow is the family of all trajectories generated by the ODE. If we denote the flow by $\\psi_t$, then\n$$ \\psi : \\mathbb{R}^d \\times [0,1] \\to \\mathbb{R}^d, \\qquad (x_0,t) \\mapsto \\psi_t(x_0) \\tag{2a} $$ For a given initial condition $X_0 = x_0$, a trajectory of the ODE is recovered via $$ X_t = \\psi_t(X_0), $$ meaning that the flow maps the initial point $x_0$ to its position at time $t$. Equivalently, the flow satisfies the ODE\n$$ \\frac{d}{dt}\\psi_t(x_0) = u_t(\\psi_t(x_0)) \\tag{2b} $$with initial condition\n$$ \\psi_0(x_0) = x_0 \\tag{2c} $$This makes the relationship between flow and trajectory very clear: a trajectory is the motion of one initial point, while the flow is the global map that moves every initial point. In other words, each trajectory is one slice of the full flow.\nThis is exactly the viewpoint needed for generative modeling. We do not want to move only one point; we want to move an entire distribution. If samples from $p_{\\text{init}}$ are evolved through the flow, then the final positions of those particles define the generated distribution.\nThe following is the flow existence and uniqueness theorem from [1]. It states that if\nTheorem 3 (Flow existence and uniqueness)\nIf $ u:\\mathbb{R}^d \\times [0,1] \\to \\mathbb{R}^d $ is continuously differentiable with a bounded derivative, then the ODE in (2) has a unique solution given by a flow $\\psi_t$. Moreover, $\\psi_t$ is a diffeomorphism for every $t$, meaning that the flow is continuously differentiable and has a continuously differentiable inverse $\\psi_t^{-1}$.\nThis is the precise mathematical statement behind the idea that each initial point is transported in one well-defined way. In machine learning, this theorem essentially always applies, because we parameterize $u_t(x)$ with neural networks and treat the required bounded-derivative regularity as part of the modeling setup. Conceptually, this gives the one-to-one correspondence we need between initial points and their transported endpoints. That existence-and-uniqueness result is what makes the transformation a valid generative mechanism rather than an ambiguous motion rule.\nHowever, there is still a practical problem. Even when the vector field $u_t$ is known, the flow $\\psi_t$ usually cannot be computed explicitly. For nontrivial vector fields, we do not have a closed-form formula for the solution of the ODE. This is the motivation for introducing a numerical method: to generate samples, we must simulate the ODE rather than solve it analytically.\nThe simplest such method is the Euler method. If we divide time into small steps of size $h$, Euler\u0026rsquo;s method replaces the continuous dynamics by the discrete update\n$$ X_{t+h} \\approx X_t + h\\,u_t(X_t). $$The idea is simple: at the current point, read off the velocity from the vector field, take a small step in that direction, and repeat. With many small steps, the discrete trajectory approximates the true continuous trajectory of the ODE. This is why Euler\u0026rsquo;s method is important for generation: it provides a concrete computational procedure for simulating the transformation from noise to data.\nA flow model uses exactly this mechanism. Instead of hand-designing the vector field, we learn a parameterized vector field $u_t^\\theta(x)$, typically with a neural network, and define the generative dynamics by\n$$ \\frac{d}{dt}X_t = u_t^\\theta(X_t), \\qquad X_0 \\sim p_{\\text{init}}. $$The model is trained so that the terminal state satisfies\n$$ X_1 \\sim p_{\\text{data}}. $$Once training is complete, sampling from the flow model is straightforward:\nSample an initial point $X_0 \\sim p_{\\text{init}}$. Choose a time discretization with step size $h$. Repeatedly apply the Euler update $$ X_{t+h} = X_t + h\\,u_t^\\theta(X_t). $$ Continue from $t=0$ to $t=1$ and return the final point $X_1$. So the whole story fits together naturally. Generation can be understood as sampling plus transformation. That transformation can be represented by an ODE, whose induced map is called a flow. Because the flow is usually not available in closed form, we simulate it numerically with Euler\u0026rsquo;s method. A trained flow model then generates data by starting from simple noise and repeatedly applying Euler steps along the learned vector field until the sample reaches the data distribution.\nReferences [1] Peter Holderrieth and Ezra Erives. Introduction to Flow Matching and Diffusion Models. 2025. https://diffusion.csail.mit.edu/\n[2] Bai-YunHan. Companion code for Introduction to Flow Matching Model. GitHub repository. https://github.com/Bai-YunHan/Companion-code-for-Introduction-to-Flow-Matching-Model\n","permalink":"https://bai-yunhan.github.io/posts/flow-matching-section-1-the-machinery-of-generation/","summary":"\u003ch1 id=\"the-machinery-of-generation\"\u003eThe Machinery of Generation\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\u003cbr\u003e\nThis 3-part introduction to flow matching is largely inspired by reference [1].\u003cbr\u003e\nIf you do not have time to complete the full course and companion code, this series is a fast, practical alternative: it distills the core ideas with high-quality visualizations and cleaner code to make the material easier to absorb.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"1-generation-as-sampling-and-transformation\"\u003e1. Generation as Sampling and Transformation\u003c/h2\u003e\n\u003cp\u003eThe starting point of generative modeling is the sampling problem. We are given training samples\u003c/p\u003e","title":"Introduction to Flow Matching Model, Part 1 of 3: The Machinery of Generation"},{"content":"The goal In the previous section, generation was framed as a transport problem: start from a simple source distribution and move samples through an ODE until they match the data distribution. In flow matching, that transport is governed by a time-dependent vector field. The central question of this section is therefore: what vector field should we use as the training target?\nFor the running toy example, the source distribution is\n$$ p_{\\text{simple}} = \\mathcal{N}(0, I_2), $$and the target distribution $p_{\\text{data}}$ is a symmetric five-mode Gaussian mixture in $\\mathbb{R}^2$. Our goal is to learn a vector field $u_t^\\theta(x)$ such that samples drawn from $p_{\\text{simple}}$ are transported to $p_{\\text{data}}$ by the ODE\n$$ \\frac{d}{dt}X_t = u_t^\\theta(X_t). $$ Figure 1. The toy setup used throughout this section. Left: the source density $p_{\\text{simple}}$ (red) and target density $p_{\\text{data}}$ (blue). Middle: trajectories induced by the learned marginal ODE. Right: the learned vector field, with colour indicating velocity magnitude. The left panel states the learning problem. The red Gaussian is easy to sample from, but it is not the distribution we want. The blue multimodal density is the data distribution we ultimately want to generate from. The model must learn how to continuously reshape the red cloud into the blue one.\nThe middle panel shows the global effect of that learned motion rule. Each black curve is one trajectory of the ODE. A single trajectory tells us how one initial sample moves over time; the whole family of trajectories is the flow that transports the source distribution toward the data distribution.\nThe right panel shows the local object we actually train: the vector field. At every point in space and time, $u_t^\\theta(x)$ tells a particle which direction to move and how fast. So if generation is the global transport, the vector field is the local rule that produces it. Constructing the correct training target for this vector field is the task of this section.\nDefinition of Probability Path A probability path is a family of distributions $\\{p_t\\}_{t \\in [0,1]}$ indexed by time, where each value of $t$ gives one intermediate distribution between the source and the target:\n$$ p_0 = p_{\\text{simple}}, \\qquad p_1 = p_{\\text{data}}. $$Instead of specifying only the start and end distributions, a probability path describes the entire transport process across time. You can think of it as a movie of distributions: $p_0$ is the first frame, $p_1$ is the last frame, and the intermediate $p_t$\u0026rsquo;s describe how the probability mass moves in between.\nFigure 2 is meant to give a qualitative feeling for this idea before we introduce more notation. In both panels, the colors mark snapshots at different times: blue for $t=0.00$, orange for $t=0.33$, green for $t=0.67$, and red for $t=1.00$. Reading the colors from blue to red, we see a distribution gradually change shape over time rather than jump directly from source to target.\nFigure 2. Two visual examples of a probability path. Left: the motion is organized around one fixed target point $z$ (red star). Right: many such transports are averaged together, producing a path for the overall data distribution. The left panel is the more local view. We pick one target point $z$, marked by the red star, and watch how the distribution moves toward that single destination. This gives a conditional probability path, written $p_t(\\cdot \\mid z)$. It is still a probability path in the same basic sense as above, but now the whole story is tied to one chosen endpoint.\nThe right panel is the more global view. Instead of focusing on one destination, we average over many target points drawn from the data distribution. The resulting family of distributions is the marginal probability path, written $p_t(\\cdot)$. So the conditional path tells a pointwise transport story, while the marginal path tells the distribution-level story we ultimately care about.\nFormally, the marginal path is obtained by averaging conditional paths over the data distribution:\n$$ p_t(x) = \\int p_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz. $$This marginal path is the distribution-level object that matters for generation, because the model only receives $(x,t)$ as input and never observes the hidden variable $z$. In practice, this creates an important asymmetry: sampling from the marginal probability path is often easy, but evaluating the marginal density is not. Formally, we can sample from $p_t$ by first drawing a data point and then sampling from the corresponding conditional path:\n$$ z \\sim p_{\\text{data}}, \\qquad x \\sim p_t(\\cdot \\mid z) \\quad \\Longrightarrow \\quad x \\sim p_t. $$But computing the density value $p_t(x)$ usually still requires the integral above.\nNote that $p_t(\\cdot \\mid z)$ and $p_t(\\cdot)$ denote distributions, while $p_t(x \\mid z)$ and $p_t(x)$ denote densities evaluated at a point $x$. In practice people often blur this distinction in notation, but conceptually it is important to keep them separate.\nWith this picture in place, we can proceed in the natural order: first define a simple conditional probability path, then understand how it induces the marginal path whose vector field we actually want to learn.\nGaussian Conditional Probability Path A particularly useful choice is the Gaussian conditional probability path\n$$ p_t(x \\mid z) = \\mathcal{N}(x;\\, \\alpha_t z,\\; \\beta_t^2 I_d), $$where the schedules $\\alpha_t$ and $\\beta_t$ satisfy\n$$ \\alpha_0 = 0,\\quad \\alpha_1 = 1,\\quad \\beta_0 = 1,\\quad \\beta_1 = 0. $$This path starts from a standard Gaussian at $t=0$ and ends at a point mass at $z$ when $t=1$. In other words, it describes how noise is gradually concentrated onto a target data point.\nThe key reason this choice is so useful is that it admits a simple reparameterization:\n$$ x_t = \\alpha_t z + \\beta_t \\epsilon, \\qquad \\epsilon \\sim \\mathcal{N}(0, I_d). $$So sampling from the conditional path is easy: scale the target point by $\\alpha_t$, scale Gaussian noise by $\\beta_t$, and add them together.\nEven better, this path gives an analytic conditional vector field. Differentiating the reparameterization with respect to time yields\n$$ \\dot{x}_t = \\dot{\\alpha}_t z + \\dot{\\beta}_t \\epsilon. $$Using $\\epsilon = \\frac{x - \\alpha_t z}{\\beta_t}$, we obtain\n$$ u_t^{\\text{target}}(x \\mid z) = \\dot{\\alpha}_t z + \\dot{\\beta}_t \\left(\\frac{x - \\alpha_t z}{\\beta_t}\\right) = \\left(\\dot{\\alpha}_t - \\frac{\\dot{\\beta}_t}{\\beta_t}\\alpha_t\\right) z + \\frac{\\dot{\\beta}_t}{\\beta_t} x. $$This is the first major win: for every triplet $(x,t,z)$ sampled from the conditional path, we can evaluate the exact target velocity $u_t^{\\text{target}}(x \\mid z)$ in closed form. So the conditional problem is tractable.\nReparameterization Trick / Change of Variable Trick The change-of-variable step above is simple but important, so it is worth writing out explicitly:\n$$\\begin{align*} (i) \\Leftrightarrow \u0026\u0026 \\dot{\\alpha}_t z + \\dot{\\beta}_t x \u0026= u_t^{\\text{target}}(\\alpha_t z + \\beta_t x | z) \\quad \\text{for all } x, z \\in \\mathbb{R}^d \\\\ (ii) \\Leftrightarrow \u0026\u0026 \\dot{\\alpha}_t z + \\dot{\\beta}_t \\left(\\frac{x - \\alpha_t z}{\\beta_t}\\right) \u0026= u_t^{\\text{target}}(x | z) \\quad \\text{for all } x, z \\in \\mathbb{R}^d \\end{align*}$$Define a new variable\n$$ x' = \\alpha_t z + \\beta_t x $$Then solve for the old $x$:\n$$ x = \\frac{x' - \\alpha_t z}{\\beta_t} $$Substitute this back into the left-hand side:\n$$ \\dot{\\alpha}_t z + \\dot{\\beta}_t x = \\dot{\\alpha}_t z + \\dot{\\beta}_t \\left(\\frac{x' - \\alpha_t z}{\\beta_t}\\right) $$At the same time, the right-hand side becomes\n$$ u_t^{\\mathrm{target}}(x' \\mid z) $$So the equation becomes\n$$ \\dot{\\alpha}_t z + \\dot{\\beta}_t \\left(\\frac{x' - \\alpha_t z}{\\beta_t}\\right) = u_t^{\\mathrm{target}}(x' \\mid z) $$Finally, since $x'$ is just a dummy variable name, rename it back to $x$. Then you get\n$$ \\dot{\\alpha}_t z + \\dot{\\beta}_t \\left(\\frac{x - \\alpha_t z}{\\beta_t}\\right) = u_t^{\\mathrm{target}}(x \\mid z) $$The practical message is that the Gaussian conditional path is not just easy to sample from; it also gives a closed-form velocity field. That makes it an ideal source of supervision.\nIn the first toy experiment we use the linear schedules\n$$ \\alpha_t = t, \\qquad \\beta_t = 1 - t. $$With this choice, the coupled sample takes the form\n$$ x_t = t z + (1-t) x_0, $$so each conditional trajectory becomes a straight interpolation from the initial noise sample $x_0$ to the target point $z$. This is why the conditional trajectories in the next figure appear visually straight.\nFigure 3. Linear beta schedule: $\\alpha_t = t$ and $\\beta_t = 1-t$. For a fixed conditioning point $z$ (red star), integrating the analytic conditional ODE reproduces the same intermediate distributions as direct sampling from the Gaussian conditional probability path. If instead we keep the same mean schedule but use the square-root beta schedule from the implementation,\n$$ \\alpha_t = t, \\qquad \\beta_t = \\sqrt{1-t}, $$then the conditional path still moves toward the same target point $z$, but the spread contracts according to a different time schedule.\nFigure 4. Square-root beta schedule: $\\alpha_t = t$ and $\\beta_t = \\sqrt{1-t}$. For a fixed conditioning point $z$ (red star), integrating the analytic conditional ODE reproduces the same intermediate distributions as direct sampling from the Gaussian conditional probability path. Figures 3 and 4 make the conditional story concrete. In the left panel, the red star is a fixed conditioning data point $z$. The black curves are trajectories generated by integrating the conditional ODE\n$$ \\frac{d}{dt}X_t = u_t(X_t \\mid z). $$Because the conditional vector field is known analytically, these trajectories are not learned approximations; they are generated from the exact conditional dynamics.\nThe middle panel shows samples from that same conditional ODE at several time slices. We begin with samples from $p_{\\text{simple}}$, integrate the ODE numerically, and record the particle locations at a few values of $t$.\nThe right panel shows direct samples from the ground-truth conditional path $p_t(x \\mid z)$ at those same time slices. Here no ODE integration is needed: we sample directly from the Gaussian distribution $\\mathcal{N}(\\alpha_t z, \\beta_t^2 I)$.\nThe key observation is that the middle and right panels match in both figures. So the conditional ODE and the conditional probability path describe the same evolving distribution under either schedule. This gives us confidence that the conditional vector field really is the correct transport rule for the chosen path.\nFrom Conditional to Marginal Probability Path The conditional story is tractable, but it is not yet the object we ultimately want to learn. A flow model takes only $(x,t)$ as input, so the appropriate target must also depend only on $(x,t)$. That means the desired object is the marginal vector field associated with the marginal path\n$$ p_t(x) = \\int p_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz. $$This is where the difficulty appears.\nSampling from the marginal path is easy: first draw $z \\sim p_{\\text{data}}$, then draw $x \\sim p_t(x \\mid z)$. But evaluating the marginal density $p_t(x)$ is generally intractable, because it requires integrating over all possible data points $z$. For the same reason, the marginal vector field is not available in closed form.\nAs stated in Theorem 10 (\u0026ldquo;Marginalization trick\u0026rdquo;) from [1]:\nTheorem 10 (Marginalization trick) For every data point $z \\in \\mathbb{R}^d$, let $u_t^{\\text{target}}(x \\mid z)$ denote a conditional vector field, defined so that the corresponding ODE yields the conditional probability path $p_t(\\cdot \\mid z)$:\n$$ X_0 \\sim p_{\\text{simple}}, \\qquad \\frac{d}{dt}X_t = u_t^{\\text{target}}(X_t \\mid z) \\quad \\Longrightarrow \\quad X_t \\sim p_t(\\cdot \\mid z). \\tag{18} $$ Then the marginal vector field $u_t^{\\text{target}}(x)$, defined by\n$$ u_t^{\\text{target}}(x) = \\int u_t^{\\text{target}}(x \\mid z)\\, \\frac{p_t(x \\mid z)\\, p_{\\text{data}}(z)}{p_t(x)}\\,dz \\tag{19} $$ follows the marginal probability path, i.e.\n$$ X_0 \\sim p_{\\text{simple}}, \\qquad \\frac{d}{dt}X_t = u_t^{\\text{target}}(X_t) \\quad \\Longrightarrow \\quad X_t \\sim p_t. \\tag{20} $$By Bayes\u0026rsquo; rule,\n$$ \\frac{p_t(x \\mid z)\\, p_{\\text{data}}(z)}{p_t(x)}=p(z \\mid x, t), $$so the same formula can also be written as\n$$ u_t^{\\text{target}}(x)=\\mathbb{E}_{z \\sim p(z \\mid x, t)}\\left[u_t^{\\text{target}}(x \\mid z)\\right]. $$$p(z\\mid x,t)$ has a simple interpretation: if we observe a particle currently at location $x$ at time $t$, how plausible is each target data point $z$ as the hidden destination that produced it? The correct marginal velocity at $(x,t)$ is obtained by averaging all conditional velocities, weighted by this posterior plausibility.\nThis is the key payoff: although we cannot directly write down the marginal vector field from the marginal density, we can construct it indirectly from conditional vector fields that are easy to analyze. In other words, conditional paths are the tractable local pieces, and the marginalization trick assembles them into the global transport rule we actually need.\nContinuity equation intuition The deeper reason why Equation (19) holds comes from the continuity equation, as stated in Theorem 12 (\u0026ldquo;Continuity Equation\u0026rdquo;) from [1]:\nTheorem 12 (Continuity Equation) Let us consider a flow model with vector field $u_t^{\\mathrm{target}}$ with $X_0 \\sim p_{\\mathrm{init}}$. Then $X_t \\sim p_t$ for all $0 \\le t \\le 1$ if and only if\n$$ \\partial_t p_t(x) = -\\operatorname{div}\\bigl(p_t u_t^{\\mathrm{target}}\\bigr)(x) \\quad \\text{for all } x \\in \\mathbb{R}^d,\\ 0 \\le t \\le 1, \\tag{24} $$ where $\\partial_t p_t(x) = \\frac{d}{dt}p_t(x)$ denotes the time-derivative of $p_t(x)$. Equation 24 is known as the continuity equation.\nEquation (24) says that the density changes only because probability mass flows through space.\nHere $u_t(x)$ is a vector field, while $p_t(x)$ is a scalar density. Their product\n$$ p_t(x)u_t(x) $$is therefore a vector field called the probability flux. It tells us how much probability mass is moving, and in which direction.\nThe divergence\n$$ \\operatorname{div}(p_tu_t)(x) $$measures the net outward flow near $x$.\nIf more probability flows out of a neighbourhood than into it, the density there decreases. If more probability flows in than out, the density there increases. So the continuity equation is the distribution-level law that links a vector field to the probability path it transports.\nFrom the continuity equation to Equation (19) Each conditional path satisfies its own continuity equation:\n$$ \\partial_t p_t(x \\mid z) = -\\operatorname{div}\\bigl(p_t(x \\mid z)\\,u_t(x \\mid z)\\bigr). $$Multiply both sides by $p_{\\text{data}}(z)$ and integrate over $z$:\n$$ \\int \\partial_t p_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz = -\\int \\operatorname{div}\\bigl(p_t(x \\mid z)\\,u_t(x \\mid z)\\bigr)\\, p_{\\text{data}}(z)\\, dz. $$Since the integral is over $z$ while the divergence acts on $x$, the two operations can be exchanged:\n$$ \\partial_t p_t(x) = -\\operatorname{div}\\left(\\int p_t(x \\mid z)\\,u_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz\\right). $$But the marginal path also satisfies a continuity equation,\n$$ \\partial_t p_t(x) = -\\operatorname{div}\\bigl(p_t(x)u_t(x)\\bigr), $$so we identify the corresponding probability flux:\n$$ p_t(x)u_t(x) = \\int p_t(x \\mid z)\\,u_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz. $$Finally, divide both sides by $p_t(x)$:\n$$ u_t(x) = \\int \\frac{p_t(x \\mid z)\\, p_{\\text{data}}(z)}{p_t(x)}\\, u_t(x \\mid z)\\, dz = \\int p(z \\mid x,t)\\, u_t(x \\mid z)\\, dz. $$This is exactly Equation (19). So the marginalization trick is not ad hoc; it is a direct consequence of how probability mass must evolve under the continuity equation.\nA Visual Demonstration of the Marginalization Trick The goal of the flow matching model is to learn the marginal vector field. In this toy experiment, we assume that the model has learned this field accurately enough that its output at the query point can be used as an approximation to the marginal vector field. Figure 5 then compares this learned vector with the posterior-weighted average of the conditional vector fields, providing a visual demonstration of Equation (19).\nFigure 5. A visual demonstration of Equation (19): at a query point $x$, the posterior-weighted average of many conditional vector fields gives the marginal vector field, and the learned model closely matches that average. The code for generating these visualizations and reproducing the experiments can be found in the companion repository [2]. Arrow What it represents Cornflower blue (fan) Individual conditional vectors $u_t(x \\mid z_i)$ from plausible sources of $x$ Crimson Monte Carlo estimate of the marginal vector in Eq. 19 Orange Learned prediction $u_t^\\theta(x)$ at the same $(x,t)$ The key comparison is between the crimson and orange arrows. The crimson arrow is not an independent ground truth; it is Eq. 19 evaluated numerically at that point. When it aligns with the orange arrow, the visualization confirms the main claim of this article: the model has learned to reproduce the posterior average of conditional fields implicitly, even though training only uses local supervision on $(z, x_t)$ pairs.\nEquation (19), derived in the previous section, is the conceptual core of Conditional Flow Matching: $$ u_t^{\\text{target}}(x) = \\mathbb{E}_{z \\sim p(z \\mid x,\\, t)}\\!\\left[\\, u_t^{\\text{target}}(x \\mid z)\\,\\right] $$The flow matching model is supposed to learn the vector field $u_t^{\\text{target}}(x)$. The difficulty is that the expectation is taken over the posterior $p(z \\mid x,t)$, so evaluating it exactly would require integrating over all of $p_{\\text{data}}$, which is intractable. Figure 5 makes this abstract identity concrete at one query point $x$: the marginal vector field is the posterior average of many conditional vector fields, and the learned model can match that average.\nHow Figure 5 Demonstrates the Marginalization Trick The goal of this section is not merely to explain how Figure 5 is drawn, but to show how the visualization is constructed to compare the learned model output with the posterior-weighted average of conditional vector fields. By unpacking each part of the construction, we gain deeper intuition for why this comparison provides a visual demonstration of Equation (19).\nTo draw this figure, we still have to answer two practical questions that Eq. 19 leaves implicit:\nWhich time $t$ should we use for the chosen query point $x$? How do we approximate the posterior expectation numerically? The rest of this section explains those two steps.\nFrom Equation 19 to a Computable Arrow The crimson arrow is obtained by using Eq. 19 itself as a Monte Carlo recipe:\n$$ u_t^{\\text{target}}(x) \\;=\\; \\mathbb{E}_{z \\sim p(z|x,t)}\\!\\left[u_t(x\\mid z)\\right] \\;\\approx\\; \\frac{1}{M}\\sum_{i=1}^{M} u_t(x \\mid z_i), \\qquad z_i \\sim p(z \\mid x, t) $$So the crimson arrow is not an external label; it is Eq. 19 evaluated numerically at that specific point.\nWhy the Header Shows $t^* = 0.39$ This subsection explains the $t^*$ shown in the header of Figure 5. In that figure, $t^* = 0.39$.\nIn the generative process, the order is fixed: $$ z \\sim p_{\\text{data}} \\;\\longrightarrow\\; t \\sim \\text{Uniform}[0,1) \\;\\longrightarrow\\; x \\sim p_t(x \\mid z) = \\mathcal{N}(\\alpha_t z,\\, \\beta_t^2 I). $$$x$ is derived from $z$ and $t$, so once we fix a query point $x$ for visualization, we cannot choose an arbitrary time slice and expect the posterior $p(z \\mid x,t)$ to be meaningful. If $x$ has near-zero density under $p_t(x)$, the posterior becomes nearly empty, the importance weights collapse onto one proposal, and the resulting visualization is mostly noise.\nFor that reason, the figure first searches for a \u0026ldquo;natural\u0026rdquo; time slice for the chosen $x$. For each candidate $t$, it estimates $$ p_t(x) = \\int p_t(x \\mid z)\\, p_{\\text{data}}(z)\\, dz \\;\\approx\\; \\frac{1}{K} \\sum_{i=1}^{K} p_t(x \\mid z_i), \\qquad z_i \\sim p_{\\text{data}}, $$ and then selects $$ t^* = \\arg\\max_{t \\in [0,1)} \\hat{p}_t(x) $$This picks the time at which the chosen $x$ is most compatible with the probability path.\nIn the implementation, this grid search over $t \\in [0.02, 0.98]$ is performed automatically before visualization, so $t$ is not exposed as a user argument.\nApproximating the Posterior with Importance Sampling Once $t^*$ has been fixed, we still need to approximate the posterior expectation in Eq. 19. Ideally we would sample from $p(z \\mid x,t)$ directly, but this is not feasible because the normalizer $$ p_t(x) = \\int p_t(x \\mid z)\\,p_{\\text{data}}(z)\\,dz $$ is intractable.\nImportance sampling solves this by drawing proposals from $p_{\\text{data}}$ and reweighting them. For any function $f$,\n$$ \\mathbb{E}_{z \\sim p(z|x,t)}[f(z)] = \\mathbb{E}_{z \\sim p_{\\text{data}}}\\!\\left[f(z)\\cdot \\underbrace{\\frac{p(z \\mid x,t)}{p_{\\text{data}}(z)}}_{\\text{importance weight } w}\\right] $$Using Bayes\u0026rsquo; rule and the fact that the proposal distribution is exactly $p_{\\text{data}}$, the weight simplifies to\n$$ w_i = \\frac{p(z_i \\mid x,t)}{p_{\\text{data}}(z_i)} = \\frac{p_t(x \\mid z_i)\\;\\cancel{p_{\\text{data}}(z_i)}}{\\underbrace{p_t(x)}_{\\text{const. in }z_i}\\;\\cancel{p_{\\text{data}}(z_i)}} \\;\\propto\\; p_t(x \\mid z_i) = \\mathcal{N}(x;\\;\\alpha_t z_i,\\;\\beta_t^2 I) $$The factor $p_t(x)$ disappears after normalization because, once $x$ and $t$ are fixed, it is the same constant for every proposal $z_i$. This is precisely why importance sampling is useful here: the intractable normalizer never has to be computed explicitly.\nIntuitively, the weight asks \u0026ldquo;if $z_i$ were the true source point, how likely would the noisy interpolation be to land at the observed query point $x$?\u0026rdquo;\n$z_i$ near $x / \\alpha_t$ → Gaussian centered close to $x$ → high weight $z_i$ far from $x / \\alpha_t$ → near-zero weight, effectively discarded From Weighted Proposals to the Arrows and Scatter Once we have normalized weights $w_i \\propto p_t(x \\mid z_i)$, we use them in two different ways.\nTo estimate the marginal vector field at the fixed location $x$, we evaluate the conditional field for every proposal $z_i$ and take the weighted average: $$ \\hat{u}_t(x) = \\sum_{i=1}^{N} w_i \\cdot u_t(x \\mid z_i), \\qquad w_i = \\text{softmax}(\\log p_t(x \\mid z_i))_i $$ This directly approximates $$ \\mathbb{E}_{z \\sim p(z \\mid x,t)}[u_t(x \\mid z)]. $$ So the crimson arrow is the posterior-weighted average of all the conditional arrows.\nFor the Panel 1 scatter plot, the goal is different. We do not want a single averaged vector; we want a cloud of points that visually behaves like samples from $p(z \\mid x,t)$. To get that, we resample $M \\ll N$ indices from the proposals according to the weights and then work with the resampled points: $$ \\hat{u}_t(x) = \\frac{1}{M} \\sum_{i=1}^{M} u_t(x \\mid \\tilde{z}_i), \\qquad \\tilde{z}_i \\sim \\text{Multinomial}(\\{z_i\\}, \\{w_i\\}) $$ For the visualization, we simply scatter the resampled points $\\tilde{z}_i$ directly.\nThese two uses of the same weights serve different purposes. The weighted average is better for estimating the vector because it keeps the exact contribution of every proposal. Resampling would throw away some of that information and add extra sampling noise. For visualization, however, resampling is exactly what we want: it converts weighted proposals into an ordinary point cloud that behaves approximately like posterior samples. This is why Figure 5 uses a weighted average for the crimson arrow but resampled points for the scatter plot. Since the conditional field is available in closed form, evaluating it on all $N = 30{,}000$ proposals is essentially free.\nSummary In the previous article, The Machinery of Generation, we explored the fundamental mechanics of generative models, uncovering the key insight that generation is essentially a process of sampling followed by transformation—specifically, using Ordinary Differential Equations (ODEs) and their simulation as the mathematical tool for this transformation. Building on that foundation, this article focused on defining the ground truth for our models—including its mathematical proofs—and detailing how to actually construct it.\nTo achieve this, we introduced the concept of a probability path to describe the transformation from the initial distribution ($p_{\\text{init}}$) to the data distribution ($p_{\\text{data}}$). We then analytically used this path to derive the ground truth vector field needed for the ODE simulation. We demonstrated that the true target for our training is the conditional vector field, rather than the marginal one, and explained how the marginal vector field can be estimated through marginalization. We supported this by analytically proving the relationship between conditional and marginal vector fields using the Continuity Equation, a fundamental result from physics. Finally, we showed qualitatively that even though the model is trained exclusively on the conditional vector field, it implicitly learns to perform this marginalization trick.\nReferences [1] Peter Holderrieth and Ezra Erives. Introduction to Flow Matching and Diffusion Models. 2025. https://diffusion.csail.mit.edu/\n[2] Bai-YunHan. Companion code for Introduction to Flow Matching Model. GitHub repository. https://github.com/Bai-YunHan/Companion-code-for-Introduction-to-Flow-Matching-Model\n","permalink":"https://bai-yunhan.github.io/posts/flow-matching-section-2-constructing-training-target/","summary":"\u003ch1 id=\"the-goal\"\u003eThe goal\u003c/h1\u003e\n\u003cp\u003eIn the previous section, generation was framed as a transport problem: start from a simple source distribution and move samples through an ODE until they match the data distribution. In flow matching, that transport is governed by a time-dependent vector field. The central question of this section is therefore: \u003cstrong\u003ewhat vector field should we use as the training target?\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor the running toy example, the source distribution is\u003c/p\u003e","title":"Introduction to Flow Matching Model, Part 2 of 3: Constructing the Training Target"},{"content":"Constructing the Training Loss Section 1 framed generative modeling as sampling through transformation: start from a simple distribution and transport samples toward the data distribution by integrating an ODE. In that picture, the learned vector field is the local motion rule, and the induced flow is the global transformation that turns noise into data.\nSection 2 then showed how to construct the relevant target vector field. By introducing conditional and marginal probability paths, we obtained a closed-form expression for the tractable conditional vector field $u_t^{\\text{target}}(x \\mid z)$ and saw, via the marginalization trick, how these conditional fields assemble into the marginal vector field $u_t^{\\text{target}}(x)$ that actually governs the distribution-level transport.\nThis leaves the next question: even if we know which vector field we want in principle, how do we train a neural network to learn it in practice? That is the problem of this section.\nThe Problem As we learned, we want the neural network $u_t^\\theta$ to approximate the marginal vector field $u_t^{\\text{target}}$. A natural way to achieve this is to use a mean-squared-error objective, namely the flow matching loss defined as\nHere, $\\text{Unif} = \\text{Unif}_{[0,1]}$ denotes the uniform distribution on $[0,1]$, and $\\mathbb{E}$ denotes expectation.\n$$\\mathcal{L}_{\\text{FM}}(\\theta) = \\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t}[\\|u_t^\\theta(x) - u_t^{\\text{target}}(x)\\|^2] \\tag{1}$$$$\\overset{(i)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)}[\\|u_t^\\theta(x) - u_t^{\\text{target}}(x)\\|^2]$$where $p_t(x) = \\int p_t(x|z) p_{\\text{data}}(z) \\mathrm{d}z$ is the marginal probability path and in $(i)$ we used the following sampling from marginal path procedure: $$ z \\sim p_{\\text{data}}, \\qquad x \\sim p_t(\\cdot \\mid z) \\qquad \\Longrightarrow \\qquad x \\sim p_t. $$ Intuitively, this loss says: First, draw a random time $t \\in [0,1]$. Second, draw a random point $z$ from our data set, sample $x \\sim p_t(\\cdot|z)$ (e.g., by adding some noise), and compute $u_t^\\theta(x)$. Finally, compute the mean-squared error between the output of our neural network and the marginal vector field $u_t^{\\text{target}}(x)$. However, we are still not done. Although Theorem 10 in [1] gives the marginalization-trick formula for $u_t^{\\text{target}}$,\n$$u_t^{\\text{target}}(x) = \\int u_t^{\\text{target}}(x|z) \\frac{p_t(x|z) p_{\\text{data}}(z)}{p_t(x)} \\mathrm{d}z, \\tag{2}$$we cannot compute it efficiently because the above integral is intractable. Instead, we will exploit the fact that the conditional velocity field $u_t^{\\text{target}}(x|z)$ is tractable. This leads to the conditional flow matching loss\n$$\\mathcal{L}_{\\text{CFM}}(\\theta) = \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)}[\\|u_t^\\theta(x) - u_t^{\\text{target}}(x|z)\\|^2]$$Note the difference to $\\mathcal{L}_{\\text{FM}}(\\theta)$ in eq. (1): we use the conditional vector field $u_t^{\\text{target}}(x|z)$ instead of the marginal vector field $u_t^{\\text{target}}(x)$. As we have an analytical formula for $u_t^{\\text{target}}(x|z)$, we can minimize the above loss easily.\nBut wait, what sense does it make to regress against the conditional vector field if it\u0026rsquo;s the marginal vector field we care about? As it turns out, by explicitly regressing against the tractable, conditional vector field, we are implicitly regressing against the intractable, marginal vector field. The next result makes this intuition precise.\nFigure 1. Comparison of Learned Marginal Probability Path and Ground Truth. Panel 1 (Trajectories of Learned Marginal ODE): We use the learned marginal vector field $u_t^\\theta$ to plot the trajectory using ODE simulation. Starting from initial samples $x_0 \\sim p_{\\text{simple}}$, we integrate the learned vector field over time using an ODE solver (e.g., Euler method). Panel 2 (Samples from Learned Marginal ODE): The samples are generated using the trajectory by extracting the points along the simulated ODE trajectories from Panel 1 at specific time steps $t$. Note that the samples appear denser here than the trajectories in Panel 1 because only a limited number of trajectories are drawn in Panel 1 for cleaner visualization. Panel 3 (Ground-Truth Marginal Probability Path): Even though we cannot compute the marginal density $p_t(x)$ directly because the integral is intractable, we can still get the ground truth sample using the marginalization trick: $$ z \\sim p_{\\text{data}}, \\qquad x \\sim p_t(\\cdot \\mid z) \\quad \\Longrightarrow \\quad x \\sim p_t. $$ In the implementation [2], this is done by sampling a target data point $z$ and then sampling $x$ from the conditional distribution $p_t(\\cdot|z)$. Because we use a Gaussian conditional probability path, $x \\sim p_t(\\cdot|z)$ is easily computable as $x = \\alpha_t z + \\beta_t \\epsilon$ where $\\epsilon \\sim \\mathcal{N}(0, I)$. Equivalence of CFM and FM Objectives: Algebraic Proof As we can see, the samples from the learned marginal ODE visually match the ground-truth probability path. This suggests that, although we train against the tractable conditional vector field $u_t^{\\text{target}}(x|z)$, the learned model still recovers the marginal behavior governed by $u_t^{\\text{target}}(x)$.\nTheorem 18. The marginal flow matching loss equals the conditional flow matching loss up to a constant. That is,\n$$ \\mathcal{L}_{\\text{FM}}(\\theta) = \\mathcal{L}_{\\text{CFM}}(\\theta) + C, $$ where $C$ is independent of $\\theta$. Therefore, their gradients coincide:\n$$ \\nabla_\\theta \\mathcal{L}_{\\text{FM}}(\\theta) = \\nabla_\\theta \\mathcal{L}_{\\text{CFM}}(\\theta). $$ Hence, minimizing $\\mathcal{L}_{\\text{CFM}}(\\theta)$ with stochastic gradient descent (SGD) is equivalent to minimizing $\\mathcal{L}_{\\text{FM}}(\\theta)$ in the same fashion. In particular, for the minimizer $\\theta^*$ of $\\mathcal{L}_{\\text{CFM}}(\\theta)$, it will hold that $u_t^{\\theta^*} = u_t^{\\text{target}}$ (assuming an infinitely expressive parameterization). [1]\nProof. The proof works by expanding the mean-squared error into three components and removing constants. [1]\n$$ \\begin{aligned} \\mathcal{L}_{\\text{FM}}(\\theta) \u0026\\overset{(i)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t} \\left[\\left\\|u_t^\\theta(x) - u_t^{\\text{target}}(x)\\right\\|^2\\right] \\\\ \u0026\\overset{(ii)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t} \\left[\\left\\|u_t^\\theta(x)\\right\\|^2 - 2\\,u_t^\\theta(x)^\\top u_t^{\\text{target}}(x) + \\left\\|u_t^{\\text{target}}(x)\\right\\|^2\\right] \\\\ \u0026\\overset{(iii)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t}\\left[\\left\\|u_t^\\theta(x)\\right\\|^2\\right] - 2\\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t}\\left[u_t^\\theta(x)^\\top u_t^{\\text{target}}(x)\\right] + \\underbrace{\\mathbb{E}_{t \\sim \\text{Unif}_{[0,1]}, x \\sim p_t}\\left[\\left\\|u_t^{\\text{target}}(x)\\right\\|^2\\right]}_{=:C_1} \\\\ \u0026\\overset{(iv)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[\\left\\|u_t^\\theta(x)\\right\\|^2\\right] - 2\\,{\\color{green}\\boxed{\\color{white}{\\mathbb{E}_{t \\sim \\text{Unif}, x \\sim p_t}\\left[u_t^\\theta(x)^\\top u_t^{\\text{target}}(x)\\right]}}} + C_1 \\end{aligned} $$where $(i)$ holds by definition, in $(ii)$ we used the formula\n$$ \\|a-b\\|^2 = \\|a\\|^2 - 2a^\\top b + \\|b\\|^2, $$in $(iii)$ we define a constant $C_1$, and in $(iv)$ we used the sampling procedure of $p_t$ given by eq. $(13)$.\n$$\\underbrace{\\mathbb{E}_{t \\sim \\text{Unif}_{[0,1]}, x \\sim p_t}\\left[\\left\\| u_t^{\\text{target}}(x) \\right\\|^2\\right]}_{=: C_1}$$ In (iii), the reason the last part is a constant is that it does not depend on the model parameters $\\theta$, while the first two terms involve $u_t^{\\theta}(x)$.\nLet us reexpress the second summand:\n$$ \\begin{aligned} {\\color{green}\\boxed{{\\color{white}\\mathbb{E}_{{\\color{orange}t \\sim \\text{Unif}, x \\sim p_t}}\\left[u_t^\\theta(x)^\\top\\right]}}} \\color{pink}{u_t^{\\text{target}}(x)} \u0026\\overset{(i)}{=} \\int_0^1 \\int p_t(x)\\, u_t^\\theta(x)^\\top u_t^{\\text{target}}(x)\\,\\mathrm{d}x\\,\\mathrm{d}t \\\\ \u0026\\overset{(ii)}{=} \\int_0^1 \\int \\cancel{p_t(x)}\\, u_t^\\theta(x) {\\color{red}\\boxed{{\\color{white}\\left[\\int u_t^{\\text{target}}(x|z)\\frac{p_t(x|z)p_{\\text{data}}(z)}{\\cancel{p_t(x)}}\\,\\mathrm{d}z\\right]}}} \\,\\mathrm{d}x\\,\\mathrm{d}t \\\\ \u0026\\overset{(iii)}{=} \\int_0^1 \\int \\int u_t^\\theta(x)^\\top u_t^{\\text{target}}(x|z)\\, p_t(x|z)\\, p_{\\text{data}}(z)\\,\\mathrm{d}z\\,\\mathrm{d}x\\,\\mathrm{d}t \\\\ \u0026\\overset{(iv)}{=} \\mathbb{E}_{\\color{orange}{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}}, x \\sim p_t(\\cdot|z)} \\left[ u_t^\\theta(x)^\\top \\color{pink}{u_t^{\\text{target}}(x|z)} \\right] \\end{aligned} $$where in $(i)$ we expressed the expected value as an integral, in $(ii)$ we use eq. $(2)$, in $(iii)$ we use the fact that integrals are linear, and in $(iv)$ we express the integral as an expected value.\nNote that this was really the crucial step of the proof:\nThe beginning of the equality used the marginal vector field $u^{\\text{target}}_t(x)$, while the end uses the conditional vector field $u^{\\text{target}}_t(x|z)$.\nRegarding (iii), it says integral is linear, let's get to the definition of linearity. For an operator $T$, it is **linear** if for any functions $f, g$ and scalars $a, b$, $$ T(af + bg) = aT(f) + bT(g). $$ This combines two properties:\nAdditivity: $T(f + g) = T(f) + T(g)$ Homogeneity: $T(af) = aT(f)$ We plug is into the equation for L FMto get:\n$$ \\begin{aligned} \\mathcal{L}_{\\text{FM}}(\\theta) \u0026\\overset{(i)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[\\left\\|u_t^\\theta(x)\\right\\|^2\\right] - 2\\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[u_t^\\theta(x)^\\top u_t^{\\text{target}}(x|z)\\right] + C_1 \\\\ \u0026\\overset{(ii)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[ \\left\\|u_t^\\theta(x)\\right\\|^2 - 2u_t^\\theta(x)^\\top u_t^{\\text{target}}(x|z) {\\color{cyan}+ \\left\\|u_t^{\\text{target}}(x|z)\\right\\|^2} {\\color{red}- \\left\\|u_t^{\\text{target}}(x|z)\\right\\|^2} \\right] + C_1 \\\\ \u0026\\overset{(iii)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[\\left\\|u_t^\\theta(x) - u_t^{\\text{target}}(x|z)\\right\\|^2\\right] + \\underbrace{\\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim p_t(\\cdot|z)} \\left[-\\left\\|u_t^{\\text{target}}(x|z)\\right\\|^2\\right]}_{C_2} + C_1 \\\\ \u0026\\overset{(iv)}{=} \\mathcal{L}_{\\text{CFM}}(\\theta) + \\underbrace{C_2 + C_1}_{=:C}. \\end{aligned} $$where in $(i)$ we plugged in the derived equation, in $(ii)$ we added and subtracted the same value, in $(iii)$ we used the formula $\\|a-b\\|^2 = \\|a\\|^2 - 2a^\\top b + \\|b\\|^2$ again, and in $(iv)$ we defined a constant in $\\theta$. This finishes the proof.\nOnce $u_t^\\theta$ has been trained, we may simulate the flow model\n$$ \\mathrm{d}X_t = u_t^\\theta(X_t)\\,\\mathrm{d}t, \\qquad X_0 \\sim p_{\\text{init}} $$via, e.g., algorithm 1 to obtain samples $X_1 \\sim p_{\\text{data}}$. Let us now instantiate the conditional flow matching loss for the choice of Gaussian probability paths:\nWhy Conditional Training Recovers the Marginal Field: Visual View The Apparent Contradiction Training minimizes the CFM loss: $$\\mathcal{L}_{\\text{CFM}}(\\theta) = \\mathbb{E}_{t,\\, z \\sim p_{\\text{data}},\\, x \\sim p_t(x|z)} \\left[\\| u_t^\\theta(x, t) - u_t^{\\text{target}}(x|z) \\|^2 \\right]$$Each training sample provides a conditional target $u_t(x|z)$ — yet the trained model learns the marginal vector field $u_t(x) = \\mathbb{E}_{z \\sim p(z|x,t)}[u_t(x|z)]$. Why?\nTraining Data Sparsity and the Contradiction Made Visible The following visualization calls visualize_relevant_training_data, which zooms into the neighbourhood of the query point $x^*$ at the optimal time $t^*$ and displays only the training samples within a small spatial and time window. The plot simultaneously reveals two things:\n1. Training data sparsity. Only a handful of points (here, $n=39$) land near any specific $(x, t)$ location. Many fall outside the bounding box entirely. Despite this, the model converges to the correct marginal vector field because neural network generalization covers the $(x, t)$ input space smoothly.\n2. The contradiction made visible. Each colour-coded arrow is the conditional target $u_t(x_i | z_i)$ for a nearby training sample, coloured by $\\log p_t(x_i | z_i)$ (plasma colormap: bright yellow = high likelihood, dark purple = low). Because each $z_i$ is different, the arrows scatter in conflicting directions — even though the inputs $(x_i, t_i)$ are nearly identical. This is the contradiction: the model receives a single input $(x, t)$ but is asked to match many different targets simultaneously. The orange arrow — the learned $u_t^\\theta(x^*)$ — is a posterior average of conditional vector field as proved by last visualization. This is not a coincidence: since $z$ is never passed to the model, the L2 loss has no other choice but to pull $u_t^\\theta(x, t)$ toward $\\mathbb{E}_{z \\sim p(z|x,t)}[u_t(x|z)]$, which is exactly the marginal vector field $u_t^{\\text{target}}(x)$.\nFigure 2. Visualization of training data. Figure 3. Visualization of relevant training data. The Key: $z$ is Hidden from the Model The model signature is $u_t^\\theta(x, t)$ — it takes only $(x, t)$ as input, never $z$.\nDuring training, $z$ appears in the data only to compute the supervision target $u_t(x|z)$; it is discarded and never passed to the network. This creates a fundamental information bottleneck.\nQuantity Available to the model during training Available at inference $x$ Yes (input) Yes (input) $t$ Yes (input) Yes (input) $z$ No — used only to compute target, then discarded No $u_t(x\\|z)$ Yes — as the loss target No Why Hidden $z$ Forces the Model to Learn the Marginal - \u0026ldquo;Visual\u0026rdquo; For any fixed input $(x, t)$, many different $z$\u0026rsquo;s are compatible with that $x$ at that time (because $\\beta_t \u003e 0$ at intermediate times, so the Gaussian clouds overlap). The model must emit a single output vector for all those different $z$\u0026rsquo;s, but the training targets $u_t(x|z_1), u_t(x|z_2), \\ldots$ all point in different directions.\nThis is a standard regression problem with a hidden variable. The fundamental result is:\nThe minimizer of $\\mathbb{E}[(y - c)^2]$ over a constant $c$ (one that cannot depend on the hidden variable) is $c^* = \\mathbb{E}[y]$.\nHere $y = u_t(x|z)$ and $c = u_t^\\theta(x, t)$ is constant with respect to $z$. So the L2-optimal prediction is: $$u_t^{\\theta*}(x, t) = \\mathbb{E}_{z \\sim p(z|x,t)}\\left[ u_t(x|z) \\right] = u_t^{\\text{target}}(x)$$The model cannot \u0026ldquo;pick a side\u0026rdquo; when it cannot see which $z$ generated the sample. The averaging is a direct consequence of the L2 loss geometry combined with the hidden $z$.\nConnection to Theorem 18 (Algebraic View) Theorem 18 arrives at the same conclusion by expanding the CFM loss: $$\\|u_t^\\theta - u_t(x|z)\\|^2 = \\|u_t^\\theta - u_t(x)\\|^2 + 2(u_t^\\theta - u_t(x))\\cdot\\underbrace{(u_t(x) - u_t(x|z))}_{\\text{zero mean over } z|x,t} + \\underbrace{\\|u_t(x) - u_t(x|z)\\|^2}_{\\text{constant in } \\theta}$$Taking $\\mathbb{E}_{z \\sim p(z|x,t)}[\\cdot]$: the cross-term vanishes because $\\mathbb{E}[u_t(x|z) \\mid x,t] = u_t(x)$, and the last term is independent of $\\theta$. Therefore $\\nabla_\\theta \\mathcal{L}_{\\text{CFM}} = \\nabla_\\theta \\mathcal{L}_{\\text{FM}}$ — same gradients, same minimizer.\nThe two perspectives are two sides of the same coin:\nRegression view: hidden $z$ + L2 loss → forced averaging → learns marginal. Algebraic view (Theorem 18): cross-term vanishes in expectation → $\\mathcal{L}_{\\text{CFM}}$ and $\\mathcal{L}_{\\text{FM}}$ share the same minimizer. Why Sparse Training Data is Sufficient The two effects play distinct roles:\nEffect What space is covered What it explains Hidden $z$ + L2 averaging Coverage of $z$\u0026rsquo;s at a fixed $(x,t)$ What the model converges to: the marginal $u_t(x)$ Neural network generalization Coverage of the $(x,t)$ input space How efficiently training reaches that target from finite data Even though the training set is a finite sample, the MLP generalizes smoothly across $(x,t)$ space, so optimizing the loss on sampled mini-batches still drives $u_t^\\theta$ toward the true minimizer everywhere. Generalization is a prerequisite — it ensures sensible outputs across all of $(x,t)$ — but it does not explain why those outputs are specifically the posterior average of conditional vector fields. That is entirely due to the hidden $z$ mechanism.\nFlow Matching Loss for Gaussian Conditional Probability Paths Let us return to the example of Gaussian probability paths $$ p_t(\\cdot|z) = \\mathcal{N}(\\alpha_t z; \\beta_t^2 I_d), $$ where we may sample from the conditional path via [1] $$ \\epsilon \\sim \\mathcal{N}(0, I_d) \\qquad \\Longrightarrow \\qquad x_t = \\alpha_t z + \\beta_t \\epsilon \\sim \\mathcal{N}(\\alpha_t z, \\beta_t^2 I_d) = p_t(\\cdot|z). \\tag{3} $$As we derived in eq. $(21)$, the conditional vector field $u_t^{\\text{target}}(x|z)$ is given by $$ u_t^{\\text{target}}(x|z) = \\left(\\dot{\\alpha}_t - \\frac{\\dot{\\beta}_t}{\\beta_t}\\alpha_t\\right)z + \\frac{\\dot{\\beta}_t}{\\beta_t}x, $$ where $\\dot{\\alpha}_t = \\partial_t \\alpha_t$ and $\\dot{\\beta}_t = \\partial_t \\beta_t$ are the respective time derivatives. Plugging in this formula, the conditional flow matching loss reads $$ \\mathcal{L}_{\\text{CFM}}(\\theta) = \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, x \\sim \\mathcal{N}(\\alpha_t z, \\beta_t^2 I_d)} \\left[ \\left\\|u_t^\\theta(x) - \\left( \\left(\\dot{\\alpha}_t - \\frac{\\dot{\\beta}_t}{\\beta_t}\\alpha_t\\right)z + \\frac{\\dot{\\beta}_t}{\\beta_t}x \\right)\\right\\|^2 \\right] $$ $$ \\overset{(i)}{=} \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, \\epsilon \\sim \\mathcal{N}(0, I_d)} \\left[ \\left\\|u_t^\\theta(\\alpha_t z + \\beta_t \\epsilon) - (\\dot{\\alpha}_t z + \\dot{\\beta}_t \\epsilon)\\right\\|^2 \\right], $$ where in $(i)$ we plugged in eq. $(3)$ and replaced $x$ by $\\alpha_t z + \\beta_t \\epsilon$. Note the simplicity of $\\mathcal{L}_{\\text{CFM}}$: we sample a data point $z$, sample some noise $\\epsilon$, and then take a mean squared error.\nLet us make this even more concrete for the special case of $\\alpha_t = t$ and $\\beta_t = 1 - t$. The corresponding probability path $$ p_t(x|z) = \\mathcal{N}(tz, (1-t)^2) $$ is sometimes referred to as the (Gaussian) CondOT probability path. Then we have $\\dot{\\alpha}_t = 1$ and $\\dot{\\beta}_t = -1$, so that $$ \\mathcal{L}_{\\text{cfm}}(\\theta) = \\mathbb{E}_{t \\sim \\text{Unif}, z \\sim p_{\\text{data}}, \\epsilon \\sim \\mathcal{N}(0, I_d)} \\left[ \\left\\|u_t^\\theta({\\color{orange}tz + (1-t)\\epsilon}) - (z - \\epsilon)\\right\\|^2 \\right]. $$Many famous state-of-the-art models have been trained using this simple yet effective procedure, e.g. Stable Diffusion 3, Meta\u0026rsquo;s Movie Gen Video, and probably many more proprietary models. In fig. 9, we visualize it in a simple example and in algorithm 3 we summarize the training procedure. [1]\n$$ \\boxed{ \\begin{aligned} \u0026 \\textbf{Algorithm 3 } \\text{Flow Matching Training Procedure (here for Gaussian CondOT path } p_t(x|z) = \\mathcal{N}(tz,(1-t)^2)\\text{)} \\\\ \u0026 \\textbf{Require: } \\text{A dataset of samples } z \\sim p_{\\text{data}}, \\text{ neural network } u_t^\\theta. \\\\[4pt] \u0026 1: \\; \\textbf{for } \\text{each mini-batch of data } \\textbf{do} \\\\ \u0026 2: \\quad \\text{Sample a data example } z \\text{ from the dataset.} \\\\ \u0026 3: \\quad \\text{Sample a random time } t \\sim \\text{Unif}_{[0,1]}. \\\\ \u0026 4: \\quad \\text{Sample noise } \\epsilon \\sim \\mathcal{N}(0, I_d). \\\\ \u0026 5: \\quad \\text{Set } x = {\\color{orange}tz + (1-t)\\epsilon}. \\hspace{2em} \\text{(General case: } x \\sim p_t(\\cdot|z)\\text{)} \\\\ \u0026 6: \\quad \\text{Compute loss} \\\\ \u0026 \\qquad \\mathcal{L}(\\theta) = \\left\\|u_t^\\theta(x) - (z - \\epsilon)\\right\\|^2 \\hspace{2em} \\text{(General case: } \\left\\|u_t^\\theta(x) - u_t^{\\text{target}}(x|z)\\right\\|^2\\text{)} \\\\ \u0026 7: \\quad \\text{Update the model parameters } \\theta \\text{ via gradient descent on } \\mathcal{L}(\\theta). \\\\ \u0026 8: \\; \\textbf{end for} \\end{aligned} } $$Intuition Behind the Flow Matching Loss The flow matching loss is given by $$ \\mathcal{L}(\\theta) = \\left\\|u_t^{\\theta}(x) - (z - \\epsilon)\\right\\|^2, \\quad \\text{where} \\; x = t z + (1-t)\\epsilon $$For the (Gaussian) CondOT path, once both $z$ and $\\epsilon$ are fixed, the trajectory $$ x_t = t z + (1-t)\\epsilon $$ is a straight line, so its velocity is the constant vector $$ \\frac{d x_t}{dt} = z-\\epsilon. $$ Accordingly, the loss trains $u_t^\\theta(x)$ to match the full target velocity $z-\\epsilon$, including both its direction and its magnitude. Importantly, this target velocity is determined jointly by $z$ and $\\epsilon$: changing $\\epsilon$ changes the straight-line path, and therefore changes the velocity as well.\nFigure 4. Linear beta schedule: $\\alpha_t = t$ and $\\beta_t = 1-t$. For fixed $z$ and $\\epsilon$, the conditional trajectory $x_t = tz + (1-t)\\epsilon$ is a straight line with constant velocity $z-\\epsilon$, so the flow matching objective trains $u_t^\\theta(x)$ to recover this full target velocity, including both direction and magnitude. References [1] Peter Holderrieth and Ezra Erives. Introduction to Flow Matching and Diffusion Models. 2025. https://diffusion.csail.mit.edu/\n[2] Bai-YunHan. Companion code for Introduction to Flow Matching Model. GitHub repository. https://github.com/Bai-YunHan/Companion-code-for-Introduction-to-Flow-Matching-Model\n","permalink":"https://bai-yunhan.github.io/posts/flow-matching-section-3-constructing-training-loss/","summary":"\u003ch1 id=\"constructing-the-training-loss\"\u003eConstructing the Training Loss\u003c/h1\u003e\n\u003cp\u003eSection 1 framed generative modeling as sampling through transformation: start from a simple distribution and transport samples toward the data distribution by integrating an ODE. In that picture, the learned vector field is the local motion rule, and the induced flow is the global transformation that turns noise into data.\u003c/p\u003e\n\u003cp\u003eSection 2 then showed how to construct the relevant target vector field. By introducing conditional and marginal probability paths, we obtained a closed-form expression for the tractable conditional vector field $u_t^{\\text{target}}(x \\mid z)$ and saw, via the marginalization trick, how these conditional fields assemble into the marginal vector field $u_t^{\\text{target}}(x)$ that actually governs the distribution-level transport.\u003c/p\u003e","title":"Introduction to Flow Matching Model, Part 3 of 3: Constructing the Training Loss"},{"content":"VAEs are an especially important model to study if you want to understand modern generative modeling for two reasons. First, they introduced a clean probabilistic latent-variable framework for generation—showing how to learn a distribution over hidden representations and sample from it in a principled way. Second, VAEs remain central to state-of-the-art generative systems today: in many diffusion-based models (notably latent diffusion and Stable Diffusion [1]), a VAE is the component that compresses images into a latent space and decodes generated latents back into images, making high-quality generation practical and efficient.\nIn this blog, I have two goals: (1) to build intuition for variational autoencoders (VAEs), and (2) to lay out the cleanest possible pseudo-code for probabilistic generation. I present the material in a top-down way—starting with what a VAE is and how it works, then diving into why it works. The mathematical perspective is largely drawn from Stanford CS231N: Deep Learning for Computer Vision (Spring 2025), Lecture 13: Generative Models 1 [2]. Alongside the post, I provide a corresponding GitHub repository with the actual training and inference code to make the ideas concrete.\nImplementation Pseudo Code Encoder The input image dimension is [$B$, $C$, $H$, $W$]. In this project, we use the MNIST dataset, where each image is grayscale and has spatial resolution $28\\times28$. A modified ResNet-18 [3], pre-trained on ImageNet, is used as the encoder. The original classification head (fc) is replaced to produce the mean (mu) and log-variance (logvar) required for the VAE\u0026rsquo;s reparameterization trick. The flow of encoder:Image → Modified ResNet → Project → Chunk → mu, logvar $[1, 28, 28]$ $\\xrightarrow{\\displaystyle \\textsf{ResNet}}$ [1, 1024] $\\xrightarrow{\\displaystyle \\textsf{Project}}$ $[1, 256]$ $\\xrightarrow{\\displaystyle \\textsf{Chunk}}$ $([1, 128], [1, 128])$ Decoder The decoder consists of one MLP layer followed by several convolutional layers (Conv) with upsampling (Up). The flow of decoder: Project -\u0026gt; Reshape -\u0026gt; Nx Conv+Up The input to the decoder is the latent code $z$. The latent dimension is [$B$, $D_{in}$] ($D_{in}$ default to 128) First it passes through a MLP → [$B$, $D_{out}$] Reshape from [$B$, $D_{out}$] → $[B, C_{init}, H_{init}, W_{init}]$, where $H_{init}=W_{init}$ Then it passes through $N$ layers of Convolution + Up-sampling layer → $[B, 1, H_{out}, W_{out}]$, where $H_{out}=W_{out}=28$ VAE The VAE takes data from MNIST dataset then pass it through a ResNet encoder. The encoder outputs the parameters $\\mu$ and $\\sigma$. Then sample latent code $z$ using $z = \\mu +\\sigma\\cdot\\epsilon$ where $\\epsilon \\sim N(0, I)$. Then pass $z$ to the decoder. Training Run input data $x$ through encoder to get distribution over $z$. Use prior loss to enforce the encoder outputs to follow a unit Gaussian distribution (zero mean, unit variance). Sample $z$ from encoder output $q_\\phi(z\\mid x)$ (Reparameterization trick). Run $z$ through decoder to get predicted data mean (Reconstruction). Use reconstruction loss to make predicted mean match $x$ under an L2 objective. Elaboration 1. The Core Concept Unlike standard Autoencoders which map Image $\\to$ Code $\\to$ Image, a VAE maps Image $\\to$ Distribution Parameters $\\to$ Image.\nThe networks do not output probabilities directly; they output the parameters (Mean $\\mu$ and Standard Deviation $\\sigma$) of Gaussian distributions.\n2. The Encoder (Inference Model) The encoder compresses high-dimensional data into a low-dimensional latent space.\nInput: Image of shape $3 \\times H \\times W$.\nOutput: Two vectors, both of length $D$ (the latent dimension, e.g., 128).\nMean Vector ($\\mu_z$): The center of the latent distribution. Log-Variance Vector ($\\log \\sigma^2_z$): The spread of the distribution. Design Choice of using Diagonal Covariance:\nWe assume the dimensions of $z$ are statistically independent. Instead of predicting a full $D \\times D$ covariance matrix (which would handle correlations between features), we only predict the diagonal.\nIntuition: This drastically reduces parameters from quadratic ($D^2$) to linear ($D$), making the model easier to train.\n3. The Reparameterization Trick The reparameterization trick is a tactical solution to a technical problem: How do we backpropagate through a random node?\nThe Issue: Inside the network, we need to sample $z$ from the distribution $q_\\phi(z|x)$ (typically a Gaussian with mean $\\mu$ and variance $\\sigma^2$. Standard random sampling breaks the chain of derivatives needed for backpropagation. The Trick: We move the randomness to an external variable $\\epsilon$ that is independent of the model parameters. The Equation: Instead of sampling $z \\sim N(\\mu, \\sigma^2)$ directly, we calculate: $z = \\mu + \\sigma \\odot \\epsilon$ where $\\epsilon \\sim N(0, 1)$ (standard normal distribution). This allows gradients to flow through $\\mu$ and $\\sigma$ during training, making the VAE end-to-end differentiable.\n4. The Decoder (Generative Model) The decoder reconstructs the image from a sampled latent point.\nInput: A vector $z$ of length $D$ (sampled from the Encoder\u0026rsquo;s distribution).\nOutput: A tensor of shape $3 \\times H \\times W$ (same as input image).\nWhat this output represents:\nMathematically, this output is the Mean Vector ($\\mu_x$) of the pixel probability distribution.\nDesign Choice (Fixed Variance):\nWe assume the pixel distribution is a Gaussian with a fixed standard deviation of 1 ($\\sigma=1$) and spherical covariance (no correlations between pixels).\n5. Critical Intuition: Why Design it This Way? A. Why Diagonal/Fixed Covariance? (The \u0026ldquo;Unmanageable Size\u0026rdquo; Problem)\nIf the Decoder tried to learn the correlations between every pair of pixels (Full Covariance), it would need a matrix of size $(H \\cdot W)^2$. For a small image, this is millions of parameters; for large images, trillions. Solution: By assuming pixels are independent (Diagonal) or fixed (Spherical), we reduce complexity from Quadratic to Linear. B. Why NOT Predict Separate Variances for Each Pixel? (The \u0026ldquo;Cheating\u0026rdquo; Problem)\nThe Idea: Why not let the decoder output a specific variance $\\sigma_i$ for every pixel $i$?\nThe Cheating Mechanism:\nAs derived in the Appendix, the Loss function includes two competing terms:\n$$ \\text{Loss} \\approx \\underbrace{\\frac{(x - \\mu)^2}{2\\sigma^2}}_{\\text{Reconstruction Error}} + \\underbrace{\\log(\\sigma)}_{\\text{Uncertainty Penalty}} $$ How it cheats: For difficult details (like edges), the reconstruction error $(x-\\mu)^2$ is naturally high. To minimize the Loss, the model can simply predict a massive variance ($\\sigma \\to \\infty$). The Result: A huge $\\sigma$ crushes the Reconstruction Error term to near zero. It is \u0026ldquo;cheaper\u0026rdquo; for the model to admit total uncertainty than to learn the hard feature. The Fix: By forcing $\\sigma=1$, the denominator is constant. The model must minimize $(x - \\mu)^2$ to lower the loss.\nC. Why is the \u0026ldquo;Reconstructed Image\u0026rdquo; just the Mean? (The MSE Connection)\nWhen you maximize the likelihood of a Gaussian where $\\sigma$ is fixed to 1, the math simplifies to:\n$\\text{Maximize } \\log P(x|z) \\iff \\text{Minimize } (x - \\mu)^2$\nTherefore, the output $\\mu$ is the value that minimizes the Mean Squared Error (L2 Loss).\nMotivation of VAE While standard auto-encoders are effective for learning feature representations (via reconstruction), they fail as generative models because their latent space ($Z$) has no enforced structure. Starting at [51:18].\nThe Limitation of Autoencoders [51:18]: if you want to use a standard autoencoder to generate new data, you would need to throw away the encoder and sample a latent vector $Z$ to pass through the decoder. However, because the autoencoder places no constraints on the latent space, you have no idea what the distribution of valid $Z$ vectors looks like. \u0026ldquo;Kicking the Can Down the Road\u0026rdquo; [51:57]: just as we didn\u0026rsquo;t know the distribution of the original data $X$, we now don\u0026rsquo;t know the distribution of the latent vectors $Z$. Therefore, we are \u0026ldquo;stuck\u0026rdquo; because we cannot easily sample a valid code to generate a realistic image. The VAE Solution [52:14]: The motivation for the VAE is to \u0026ldquo;force some structure on the $Z$\u0026rsquo;s.\u0026rdquo; By forcing the latent space to approximate a known distribution (typically a unit Gaussian), the VAE ensures that we can easily sample a random $Z$ from that known distribution and pass it through the decoder to generate valid new data. VAEs are essentially a \u0026ldquo;probabilistic spin\u0026rdquo; on traditional autoencoders designed specifically to enable this sampling capability [53:08].\nFormulation Goal — Maximize Marginal Likelihood We want to find parameters $\\theta$ such that the data we observed ($x$) is highly probable under our model. We maximize the Marginal Likelihood, not the Conditional Likelihood.\nEquation: $\\theta^* = \\operatorname*{argmax}_\\theta \\sum_{i} \\log p_\\theta(x^{(i)})$\nWhy Marginal $p_\\theta(x)$: The \u0026ldquo;marginal\u0026rdquo; integrates over all possible latent variables $z$.\n$p(x\\mid z)$ (Single Case): This asks, \u0026ldquo;How likely is this image if the hidden concept is exactly $z$?\u0026rdquo; $p(x)$ (Collectively all cases): This asks, \u0026ldquo;How likely is this image considering every possible hidden concept that could produce it?\u0026rdquo; i.e. $p(x) = \\int p(x|z)p(z)dz$. Symbol Name Description $p_\\theta(x \\mid z)$ Likelihood How likely is the image $x$, given the hidden traits $z$? $p_\\theta(z)$ Prior What is the distribution of hidden traits $z$ before seeing data? $p_\\theta(z \\mid x)$ Posterior Given the image $x$, what are the likely traits $z$? $p_\\theta(x)$ Marginal Likelihood How probable is the data $x$ overall, summing over all possible $z$\u0026rsquo;s? Why $p_\\theta(x)$ is called marginal likelihood?\nHistorically, probability distributions were written in tables. For example, if you had a joint distribution over $x$ and $z$:\nz=0 z=1 Row total x=0 … … p(x=0) x=1 … … p(x=1) Column total p(z=0) p(z=1) 1 To get the probability of $x$ alone, you would sum across the row, and the result would appear in the margin of the table. Hence: Summing out a variable = looking at the margin of the joint distribution. That operation was literally written in the margins of the table → “marginal probability.” The process of computing it → “marginalization.”\nBlocker ← The Intractability To calculate the marginal likelihood, we face a mathematical dead end because we cannot compute the integral or the posterior.\nThe Integral Problem: $p_\\theta(x) = \\int p_\\theta(x|z)p(z) \\, dz$\nFor complex data (like images), this integral is impossible to calculate (intractable) because it requires summing over an infinite number of possible $z$ configurations.\nThe Bayes\u0026rsquo; Rule Problem:\nWe might try to find $p_\\theta(x)$ via Bayes\u0026rsquo; Rule:\n$$ p_\\theta(x) = \\frac{p_\\theta(x\\mid z)p(z)}{p_\\theta(z\\mid x)} $$However, we only have a decoder to compute $p_\\theta(x\\mid z)$. To compute the denominator $p_\\theta(z|x)$ (the true posterior), we need to use the Bayes’ rule which requires knowing $p_\\theta(x)$, $p_\\theta(z\\mid x) = \\frac{p_\\theta(x\\mid z)p(z)}{p_\\theta(x)}$, so it is intractable.\nSolution ← Variational Inference (Replacing $p_\\theta$ with $q_\\phi$) Since the true posterior $p_\\theta(z|x)$ is impossible to calculate, we approximate it with a tractable distribution $q_\\phi(z|x)$ (the Encoder/Neural Network).\nThe Approximation: $q_\\phi(z|x) \\approx p_\\theta(z|x)$ Intuition The reconstruction loss and prior loss is fighting with each other This \u0026ldquo;tug-of-war\u0026rdquo; is one of the most fascinating parts of a VAE because you can actually see the result of these two forces fighting when you plot the latent space.\nHere is the visual evidence that the KL divergence doesn\u0026rsquo;t just \u0026ldquo;force\u0026rdquo; everything to zero, but rather organizes it.\n1. The \u0026ldquo;Healthy\u0026rdquo; VAE Balance When the Reconstruction Loss (keep data distinct) and KL Loss (keep data Gaussian) are balanced correctly, the latent space looks like this:\nWhat to notice:\nGlobal Centering: Notice that the entire cloud of points is centered around $(0,0)$ and roughly spans between $-3$ and $3$ (typical for a unit variance). This is the KL loss doing its job. Local Distinctness: Crucially, individual digits are not all collapsed to $(0,0)$. The \u0026ldquo;7\u0026quot;s might be clustered at $(-1, 2)$ and the \u0026ldquo;0\u0026quot;s at $(1, -1)$. The encoder has learned to shift the mean $\\mu$ away from 0 just enough to distinguish the digits, but keeps them packed tight enough to satisfy the Gaussian prior. Smoothness: The clusters touch each other. This means if you sample a point halfway between a \u0026ldquo;1\u0026rdquo; and a \u0026ldquo;7\u0026rdquo;, you get a digit that looks like a plausible mix of both. 2. What happens without the KL \u0026ldquo;Force\u0026rdquo; (Standard Autoencoder) If you remove the KL term (essentially predicting mean/variance but with no penalty for where they are), the Reconstruction Loss takes over completely.\nWhat to notice:\nGaps and Explosions: The clusters fly apart. The model might put \u0026ldquo;0\u0026quot;s at $(100, 100)$ and \u0026ldquo;1\u0026quot;s at $(-50, -50)$ just to be absolutely sure it doesn\u0026rsquo;t confuse them. No Structure: There is no center. The spread is arbitrary. Dead Zones: There are massive empty gaps between clusters. If you try to generate an image from those gaps, you get static/noise because the decoder has never seen data there. 3. What happens if the KL Force Wins (Posterior Collapse) This is the scenario you feared—where the model is \u0026ldquo;forced\u0026rdquo; to 0 and 1.\nThe Visual: Imagine the first plot, but all the colored clusters are stacked directly on top of each other at $(0,0)$. The Result: The encoder output is always $\\mu=0, \\sigma=1$ effectively ignoring the input image. The decoder receives pure noise every time and produces a single, blurry \u0026ldquo;average\u0026rdquo; image (like a gray ghost) for every single input. 4. Summary The KL divergence acts like a spring attached to the origin $(0,0)$.\nReconstruction tries to pull the data points apart so they don\u0026rsquo;t overlap. KL Divergence (the spring) pulls them back toward the center. The final state is a tense equilibrium: data points distinct enough to be recognized, but bunched tight enough to form a smooth, continuous space. Role of $z_i$ in latent code $Z$ Auto-Encoding Variational Bayes, ICLR 2014 [4].\nIn a Variational Autoencoder (VAE), $Z$ represents the latent code—a compressed, hidden representation of the input data. $z_i$ refers to a specific individual dimension (or component) within this vector. The core idea shown here is disentanglement: the model attempts to map distinct, meaningful semantic features of the data to separate dimensions ($z_i$). For example, changing the value of one dimension ($z_1$) might smoothly transform the digit\u0026rsquo;s identity (e.g., changing a 6 to a 0), while changing another dimension ($z_2$) might strictly alter the slant or thickness of the writing, without changing the digit itself.\nImplementation of sampling $z_1$ and $z_2$ and set rest of $z_i$ to zeros.\nLimitation: Blurry results (i.e. What’s Next?) The nature of loss function (MSE/log-likelihood) tend to produce \u0026ldquo;blurry\u0026rdquo; results in VAE.\nThe \u0026ldquo;blurry VAE\u0026rdquo; phenomenon is a direct mathematical consequence of how we measure \u0026ldquo;error\u0026rdquo; using Mean Squared Error (MSE) or Log-Likelihood.\nIn short: MSE forces the model to be a conservative \u0026ldquo;average,\u0026rdquo; rather than a bold \u0026ldquo;guesser.\u0026rdquo;\nHere is the breakdown of why this happens.\n1. The \u0026ldquo;Safety\u0026rdquo; of the Average Imagine the model is trying to reconstruct a picture of a zebra, but the latent space is slightly uncertain about exactly where a specific stripe should be.\nPossibility A: The stripe is at pixel 100. Possibility B: The stripe is at pixel 101. If the model guesses A (sharp stripe at 100), but the truth was B, the MSE penalty is massive because the pixel values are totally opposite (black vs. white).\nIf the model guesses B (sharp stripe at 101), but the truth was A, the penalty is essentially double (wrong on both pixels).\nThe VAE\u0026rsquo;s Solution: It predicts a gray smear across pixels 100 and 101.\nWhy? Gray is never \u0026ldquo;perfectly right,\u0026rdquo; but it is never \u0026ldquo;catastrophically wrong.\u0026rdquo; It minimizes the squared error across all plausible possibilities. The model learns to hedge its bets to lower the loss. 2. The Unimodal Assumption (Mathematics) This is the technical root of the problem.\nMSE is equivalent to Maximum Likelihood under a Gaussian distribution. When you use MSE, you are implicitly telling the model: \u0026ldquo;Assume the pixel value comes from a single Bell curve (Unimodal).\u0026rdquo; The Problem: Real data is Multimodal.\nA pixel at the edge of an object could plausibly be Black (background) OR White (object). It is almost never Gray.\nMultimodal Reality: Two peaks (one at 0, one at 255). Unimodal Constraint: The model must fit one Bell curve to explain both peaks. Result: The model centers the Bell curve right in the middle (Gray/Blur) to cover both options. 3. High Frequency vs. Low Frequency Low Frequency (Structure): Where is the head? Where is the background? VAEs are great at this because the \u0026ldquo;average\u0026rdquo; of a head is still roughly a head shape. High Frequency (Texture/Edges): Where is this specific hair strand? Where is the pore on the skin? The \u0026ldquo;average\u0026rdquo; of many possible hair strand positions is just a smooth blur. Since VAEs optimize for the average case, they effectively apply a low-pass filter to the image, smoothing out all the sharp \u0026ldquo;high frequency\u0026rdquo; noise that our eyes interpret as realistic detail. Comparison: Why GANs [5] don\u0026rsquo;t blur VAE (MSE): \u0026ldquo;I must be close to the pixel values on average. I will be safe and blurry.\u0026rdquo; GAN (Discriminator): \u0026ldquo;I don\u0026rsquo;t care if the stripe is at pixel 100 or 101, but if it\u0026rsquo;s gray/blurry, the Discriminator will know it\u0026rsquo;s fake. I must pick one sharp location, even if I guess wrong.\u0026rdquo; Loss Function Strategy Result MSE / L2 Minimize variance; fit the mean of the distribution. Blurry. The mean of \u0026ldquo;sharp left\u0026rdquo; and \u0026ldquo;sharp right\u0026rdquo; is \u0026ldquo;blurry middle.\u0026rdquo; L1 Loss Minimize absolute error; fit the median. Slightly sharper, but still blurry compared to GANs. Adversarial Fool a judge; match the distribution. Sharp. Forces a decision (collapse to a mode) rather than an average. The Perceptual Loss (VGG Loss) [6] or VQ-VAE [7] are the standard modern techniques specifically designed to fix this blurriness in VAEs without needing a full GAN setup.\nRetrospection Question-1: In a VAE, does the KL term make each class (e.g., each digit 0–9) form its own Gaussian cluster in latent space near zero, such that the mixture of the 10 class-wise clusters matches a standard normal, and the class separability comes from different encoder-produced $\\mu$ and $\\sigma$?\nConcretely, We sample $z$ from the encoder, $z \\sim q_\\phi(z|x)$, where $q_\\phi(z|x)$ is a gaussian distribution, i.e. $z \\sim N(\\mu, \\sigma^2)$. We use KL divergence loss to make $N(\\mu, \\sigma^2)$ look similar to a standard gaussian distribution. $z$ is a latent code (vector of length $D$). Say, we are reconstructing hand written digit of 10 classes (digit 0 ~ 9). The $z$ representing digit $0$ is a gaussian distribution, and the $z$ representing digit $1$ is another gaussian distribution. These cluster both locate near to zero mean. The totality of the 10 cluster conform to the standard gaussian distribution. The distinctiveness between each cluster comes from the difference in $\\mu$ and $\\sigma$ generated by the encoder.\nAnswer: The understanding is correct.\nIn summary, the VAE is a Master of Compromise\nThe VAE does clustering (via the reconstruction term), where each class (digit) gets its own localized cluster defined by its mean $\\mu$. The VAE does regularization (via the KL term), ensuring that all these clusters stay packed together in a smooth, continuous space that collectively resembles the $N(\\mu, \\sigma^2)$ distribution. The Local View (Individual Input) For any single input (e.g., a specific image of a digit \u0026lsquo;2\u0026rsquo;), the encoder outputs a specific $\\mu$ and $\\sigma$. This defines a \u0026ldquo;neighborhood\u0026rdquo; in the latent space where that specific image lives. By sampling $z$ from this neighborhood, the decoder learns that any point in this small area should look like that specific \u0026lsquo;2\u0026rsquo;. The Class View (Group of Inputs belong to same class) As you noted, all the \u0026lsquo;2\u0026rsquo;s will have their own $\\mu$ and $\\sigma$ values. Because they all share similar visual features, the Reconstruction Loss naturally forces their $\\mu$ values to be near each other. This creates a cluster (a Gaussian Mixture component) for the digit \u0026lsquo;2\u0026rsquo;. The Global View (The Prior) The KL Divergence is the \u0026ldquo;global supervisor.\u0026rdquo; It doesn\u0026rsquo;t care about the labels (0–9); it only sees the totality of all these neighborhoods. It exerts a pull on every individual distribution to stay close to $0$ and have a spread near $1$. The result is that the \u0026ldquo;cloud of clusters\u0026rdquo; (the totality) conforms to the standard Gaussian shape. A Helpful Mental Model: \u0026ldquo;The Bubble Map\u0026rdquo; Imagine each input image is a bubble.\nThe Reconstruction Loss wants the bubbles to be solid and distinct so the decoder can \u0026ldquo;see\u0026rdquo; the image clearly. The KL Loss wants all the bubbles to move to the center of the map and be exactly the same size. The Training Process is the struggle to pack all these bubbles into a small, circular container (the Standard Normal Prior) without them overlapping so much that they lose their identity. Question-2: What is the use of the randomness in VAE ?\nConcretely, let’s say we are reconstructing hand written digit of 10 classes (digit 0 ~ 9).\nRegarding the encoder, when given 10 input images of digit ‘1’, the output $\\mu$ and $\\sigma$ would be very close if not the same. Namely, the latent code $z$ of the image of digit ‘1’ belong to the same cluster. Am i right?\nDuring inference, if we feed the same image of ‘1’ to the encoder, though the $\\mu$ and $\\sigma$ output by encoder is the same, due to the randomness of $\\epsilon$, the output $z$ would be different. namely, the latent code for the same input image is not deterministic. Am i right? In contradiction, for auto-encoder, the latent code for the same input image is deterministic. What is the use of the randomness?\nAnswer: Randomness is the heart of the VAE magic.\n1. Verification of Intuition\nPoint 1: Do 10 images of digit \u0026lsquo;1\u0026rsquo; cluster together?\nYes, you are right. The encoder will map all distinct images of \u0026lsquo;1\u0026rsquo; to values of $\\mu$ and $\\sigma$ that are close to each other in the latent space. Nuance: They won\u0026rsquo;t be exactly the same. One \u0026lsquo;1\u0026rsquo; might be slanted (mapping to slightly left in the cluster), and another might be bold (mapping to slightly right). The VAE captures these stylistic differences in the precise values of $\\mu$. Point 2: Is the latent code $z$ stochastic for the same image?\nYes, you are right. If you feed the exact same image into a VAE multiple times (and you are using the sampling step), you will get a slightly different vector $z$ each time because of the random noise $\\epsilon$. Note: In practical deployment (e.g., if using a VAE just to compress data), engineers often skip the sampling and just use $\\mu$ to get a deterministic code. But strictly speaking, the mathematical definition of the VAE inference path involves this randomness. 2. The Core Question: What is the use of the randomness?\nYou asked: \u0026ldquo;For auto-encoder, the latent code is deterministic. What is the use of the randomness?\u0026rdquo; This is the most critical concept in VAEs. The randomness transforms the latent space from a discrete lookup table into a continuous landscape. Here is the analogy: The Dot vs. The Bubble.\nA. The Standard Autoencoder (The Dot)\nMechanism: It maps an input image to a single, precise point (a dot) in space. The Problem: The decoder only learns to decode that specific point. The Consequence: If you sample a point just slightly next to that dot—in the \u0026ldquo;empty space\u0026rdquo; between a \u0026lsquo;1\u0026rsquo; and a \u0026lsquo;2\u0026rsquo;—the decoder has no idea what to do. It often produces garbage or static because it never learned to handle that specific coordinate during training. The latent space is full of \u0026ldquo;holes.\u0026rdquo; B. The VAE (The Bubble)\nMechanism: By predicting a mean $\\mu$ and variance $\\sigma$ and adding noise, the encoder maps the input image not to a dot, but to a cloud or bubble of probability. The Effect: During training, the decoder is forced to reconstruct the digit \u0026lsquo;1\u0026rsquo; not just from a single point $\\mu$, but from any point sampled within that bubble $z$. The \u0026ldquo;Use\u0026rdquo; of Randomness: Forcing Continuity (Smoothness): Because the decoder must reconstruct a \u0026lsquo;1\u0026rsquo; from anywhere inside the bubble, it learns that points near each other should produce similar outputs. This eliminates the \u0026ldquo;holes.\u0026rdquo; If two bubbles (say, a \u0026lsquo;1\u0026rsquo; and a \u0026lsquo;7\u0026rsquo;) overlap slightly, the decoder learns to generate a hybrid digit in that overlapping region. Dense Packing: The KL divergence (the \u0026ldquo;spring\u0026rdquo; we discussed earlier) tries to pack these bubbles as close to the center as possible without crushing them. Because they are bubbles (taking up volume) and not dots (infinitely small), they fill up the latent space completely. The randomness prevents the model from \u0026ldquo;memorizing\u0026rdquo; specific points. It forces the model to learn a region for each digit, ensuring that the latent space is smooth, continuous, and safe to sample from for generation.\nAppendix Derivation of the Gaussian Loss Here is the derivation of why the loss function looks the way it does, starting from the definition of the Gaussian distribution.\n1. The Gaussian Probability Density Function (PDF)\nFor a single pixel value $x$, modeled by a Gaussian with mean $\\mu$ and standard deviation $\\sigma$:\n$$ P(x|\\mu, \\sigma) = \\frac{1}{\\sqrt{2\\pi\\sigma^2}} \\cdot e^{-\\frac{(x - \\mu)^2}{2\\sigma^2}} $$2. Log-Likelihood\nTo train models, we want to maximize the probability of the true data. We take the Logarithm to make the math easier (Log is monotonic, so maximizing Log($P$) is the same as maximizing $P$).\n$$ \\log P(x) = \\log\\left(\\frac{1}{\\sqrt{2\\pi\\sigma^2}}\\right) + \\log\\left(e^{-\\frac{(x - \\mu)^2}{2\\sigma^2}}\\right) $$Using logarithm rules $\\log(e^y) = y$ and $\\log(1/a) = -\\log(a)$):\n$$ \\log P(x) = -\\frac{1}{2}\\log(2\\pi\\sigma^2) - \\frac{(x - \\mu)^2}{2\\sigma^2} $$3. Negative Log-Likelihood (The Loss)\nIn Deep Learning, we minimize Loss, which is the Negative Log-Likelihood. We flip the signs:\n$$ \\text{Loss} = \\underbrace{\\frac{1}{2}\\log(2\\pi)}_{\\text{Constant}} + \\underbrace{\\log(\\sigma)}_{\\text{Variance Term}} + \\underbrace{\\frac{(x - \\mu)^2}{2\\sigma^2}}_{\\text{Error Term}} $$4. Simplified Loss for Analysis\nIgnoring the constant (since it doesn\u0026rsquo;t change with weights), we get the equation used in the \u0026ldquo;Cheating\u0026rdquo; section:\n$$ \\text{Loss} \\propto \\frac{(x - \\mu)^2}{2\\sigma^2} + \\log(\\sigma) $$Derivation of the ELBO (Evidence Lower Bound) The lecture did a great job in explaining ELBO. [1:03:04 to 1:08:55]\n$$ \\log p_\\theta(x) = \\log \\frac{p_\\theta(x|z) p(z)}{p_\\theta(z|x)} $$Multiply top and bottom by $\\textcolor{lightblue}{q_\\phi(z|x)}$\n$$ \\log p_\\theta(x) = \\log \\frac{p_\\theta(x|z)p(z)}{p_\\theta(z|x)} = \\log \\frac{p_\\theta(x|z)p(z)\\textcolor{lightblue}{q_\\phi(z|x)}}{p_\\theta(z|x)\\textcolor{lightblue}{q_\\phi(z|x)}} $$Logarithms + rearranging:\n$$ \\begin{align*} \\log p_\\theta(x) \u0026= \\log \\frac{p_\\theta(x \\mid z)p(z)}{p_\\theta(z \\mid x)} = \\log \\frac{\\textcolor{cyan}{p_\\theta(x \\mid z)}\\textcolor{green}{p(z)}\\textcolor{red}{q_\\phi(z \\mid x)}}{\\textcolor{yellow}{p_\\theta(z \\mid x)}\\textcolor{lightblue}{q_\\phi(z \\mid x)}} \\\\[12pt] \u0026= \\log \\textcolor{cyan}{p_\\theta(x \\mid z)} - \\log \\frac{\\textcolor{lightblue}{q_\\phi(z \\mid x)}}{\\textcolor{green}{p(z)}} + \\log \\frac{\\textcolor{red}{q_\\phi(z \\mid x)}}{\\textcolor{yellow}{p_\\theta(z \\mid x)}} \\end{align*} $$[1:04:01 - 1:04:58] We can wrap in an expectation since it doesn’t depend on $z$: $\\log p_\\theta(x) = E_{z \\sim q_\\phi(z \\mid x)}\\left[\\log p_\\theta(x)\\right]$. The expectation (average) of a constant value is just the constant itself.\nIn the notation $E_{z \\sim q_\\phi(z \\mid x)}\\left[\\log p_\\theta(x)\\right]$: $z$ is indeed sampled from the distribution $q_\\phi$. It is conditioned on the input $x$ (\u0026ldquo;given x\u0026rdquo;). $q_\\phi(z \\mid x)$ is the Encoder. In plain English: \u0026ldquo;We are going to calculate the average value of [whatever term follows] by trying out many different latent codes ($z$) that our Encoder thinks are likely for this specific image ($x$).\u0026rdquo; The term they are looking at is logP(x). This represents the probability of the image (data) occurring. The variable $z$ represents the latent code (hidden features). Crucially: $\\log P(x)$ depends only on the data $x$. It does not depend on $z$. Therefore, relative to $z$, the term $\\log P(x)$ is a constant. $$ \\log p_\\theta(x) = E_z[\\log p_\\theta(x \\mid z)] - E_z \\left[ \\log \\frac{q_\\phi(z \\mid x)}{p(z)} \\right] + E_z \\left[ \\log \\frac{q_\\phi(z \\mid x)}{p_\\theta(z \\mid x)} \\right] $$The 2nd and 3rd term are KL divergence which measures dissimilarity between two probability distributions.\n$$ \\begin{align*} \\log p_{\\theta}(x) \u0026= \\log \\frac{p_{\\theta}(x \\mid z)p(z)}{p_{\\theta}(z \\mid x)} = \\log \\frac{p_{\\theta}(x \\mid z)p(z)q_{\\phi}(z \\mid x)}{p_{\\theta}(z \\mid x)q_{\\phi}(z \\mid x)} \\\\ \u0026= E_{z}\\left[\\log p_{\\theta}(x \\mid z)\\right] - E_{z}\\left[\\log \\frac{q_{\\phi}(z \\mid x)}{p(z)}\\right] + E_{z}\\left[\\log \\frac{q_{\\phi}(z \\mid x)}{p_{\\theta}(z \\mid x)}\\right] \\\\ \u0026= E_{z \\sim q_{\\phi}(z \\mid x)}\\left[\\log p_{\\theta}(x \\mid z)\\right] - D_{KL}\\left(q_{\\phi}(z \\mid x), p(z)\\right) + D_{KL}\\left(q_{\\phi}(z \\mid x), p_{\\theta}(z \\mid x)\\right) \\end{align*} $$$E_{z \\sim q_{\\phi}(z \\mid x)}\\left[\\log p_{\\theta}(x \\mid z)\\right]$: The data reconstruction term. x→encoder→decoder should reconstruct $x$. Can compute in closed form for Gaussians.\n$D_{KL}\\left(q_{\\phi}(z \\mid x), p(z)\\right)$: The prior term. We are forcing the Encoder to organize its output (i.e. the hidden codes $z$) so that, overall, they form a nice, neat Standard Gaussian distribution. Can compute in closed form for Gaussians.\nThe \u0026ldquo;Encoder Output\u0026rdquo;: This is $q_\\phi(z \\mid x)$. When you feed an image $x$ into the encoder, it doesn\u0026rsquo;t just give you a single point; it gives you a probability distribution (specifically, it predicts a mean $\\mu$ and a variance $\\sigma^2$ for that image). The \u0026ldquo;Prior\u0026rdquo;: This is $p(z)$. We assume this is a Standard Unit Gaussian (a bell curve centered at 0 with a width of 1). The \u0026ldquo;Match\u0026rdquo;: The goal of training is to minimize the difference (KL Divergence) between the two. $D_{KL}\\left(q_{\\phi}(z \\mid x), p_{\\theta}(z \\mid x)\\right)$: Posterior Approximation. Encoder output $q_\\phi(z\\mid x)$ should match $p_\\theta(z\\mid x)$. We cannot compute this for Gaussians.\nKL is ≥ 0, so we can drop it to get lower bound on likelihood. This is out VAE training objective. Jointly train encoder $q$ and decoder $p$ to maximize the variational lower bound on the data likelihood. Also called Evidence Lower Bound (ELBo)\n$$ \\log p_{\\theta}(x) \\geq E_{z \\sim q_{\\phi}(z|x)} \\left[ \\log p_{\\theta}(x|z) \\right] - D_{KL} \\left( q_{\\phi}(z|x), p(z) \\right) $$\nCitation Please cite this work as:\nBai, Yechao. \u0026#34;VAE Revisited 2026: The Foundation of Generative AI\u0026#34;. Yechao\u0026#39;s Log (Jan 2026). https://bai-yunhan.github.io/posts/vae-variational-auto-encoder Or use the BibTex citation:\n@article{Bai2026VAE, title = {VAE Revisited 2026: The Foundation of Generative AI}, author = {Bai, Yechao}, journal = {bai-yunhan.github.io}, year = {2026}, month = {Jan}, url = \u0026#34;https://bai-yunhan.github.io/posts/vae-variational-auto-encoder/\u0026#34; } References High-Resolution Image Synthesis with Latent Diffusion Models (Stable Diffusion)\nRombach, R., Blattmann, A., Lorenz, D., Esser, P., \u0026amp; Ommer, B. (2022). Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).\narXiv:2112.10752\nStanford CS231N: Deep Learning for Computer Vision (Lecture 13, Spring 2025)\nYouTube Link\nDeep Residual Learning for Image Recognition (ResNet)\nHe, K., Zhang, X., Ren, S., \u0026amp; Sun, J. (2016). Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).\narXiv:1512.03385\nAuto-Encoding Variational Bayes (Original VAE Paper)\nKingma, D. P., \u0026amp; Welling, M. (2014). Auto-Encoding Variational Bayes. International Conference on Learning Representations (ICLR).\narXiv:1312.6114\nGenerative Adversarial Nets (GAN)\nGoodfellow, I., et al. (2014). Advances in Neural Information Processing Systems (NeurIPS).\narXiv:1406.2661\nPerceptual Losses for Real-Time Style Transfer and Super-Resolution\nJohnson, J., Alahi, A., \u0026amp; Fei-Fei, L. (2016). European Conference on Computer Vision (ECCV).\narXiv:1603.08155\nNeural Discrete Representation Learning (VQ-VAE)\nvan den Oord, A., Vinyals, O., \u0026amp; Kavukcuoglu, K. (2017). Advances in Neural Information Processing Systems (NeurIPS).\narXiv:1711.00937\n","permalink":"https://bai-yunhan.github.io/posts/vae-variational-auto-encoder/","summary":"\u003cp\u003eVAEs are an especially important model to study if you want to understand modern generative modeling for two reasons. First, they introduced a clean probabilistic latent-variable framework for generation—showing how to learn a distribution over hidden representations and sample from it in a principled way. Second, VAEs remain central to state-of-the-art generative systems today: in many diffusion-based models (notably latent diffusion and Stable Diffusion [1]), a VAE is the component that compresses images into a latent space and decodes generated latents back into images, making high-quality generation practical and efficient.\u003c/p\u003e","title":"VAE Revisited 2026: The Foundation of Generative AI"}]