Adversarial

Implicit generative models as factors — models you can sample from but whose density you cannot evaluate — and the discriminator read as what it is: a density-ratio estimator.

Three factors

src = NoiseSource(:z, 8; nsamples = 512)              # q(z), emits particles
gen = GeneratorFactor(my_decoder, (z = 8, x = 2))     # x = G_θ(z)
rat = RatioFactor(my_critic, 2)                        # log p(x) - log q(x)

All three are unidirectional. A generator is a function with no residual behind it, so there is nothing to run backwards — inverting $G_\theta$ is the GAN-inversion problem, and factor_message on the latent channel says so rather than pretending.

This package is the only one that produces a SampleBelief: a weighted particle set rather than a point or a parametric distribution.

The ratio identity

Train a classifier to separate $p$ from $q$ with equal class priors, and at the optimum

\[\operatorname{logit} D^\ast(x) = \log p(x) - \log q(x)\]

so the discriminator's logit is the log density ratio. logratio and discriminator are the two readings of that number.

Reweighting

The reason a ratio is useful here has little to do with generating images: importance reweighting needs a ratio, not a density, and that is what a classifier gives you.

pb, st = reweight(rat, particle_belief, ps, st)   # w_i ∝ w_i · r(x_i)
effective_sample_size(pb)                          # ...and check it meant something
weighted_mean(pb), weighted_cov(pb)

Always look at effective_sample_size. Reweighting concentrates, and the particle set can collapse onto a handful of samples with no error raised.

What it cannot do

Train adversarially. The generator descends the same quantity the discriminator ascends, and the graph's free energy has one sign. You can express a GAN's wiring here — and it is a DAG, since every factor is unidirectional — but not its objective. local_free_energy on a RatioFactor computes the number both players care about; nothing acts on it.

Also: no training of either network, and reweight is not idempotent (applying the same ratio twice squares the weights), so it is not wired into Mycelium.combine.

API

Adversarial.Adversarial — Module
Adversarial

Implicit generative models as Lenticulum factors — Mohamed & Lakshminarayanan's Learning in Implicit Generative Models (arXiv:1610.03483), and the GAN diagram read as a factor graph.

The three factors

factorispolarities
NoiseSourcethe latent prior $q(z)$, as an emitting factor1
GeneratorFactor$x = G_\theta(z)$ — samples out, no density1
RatioFactor$\log r(x) = \log p(x) - \log q(x)$ — the discriminator1

Every one of them is unidirectional, and that is the finding rather than a shortcoming: a generator is a function with no residual behind it, so there is nothing to run backwards. See [[Three Senses of Implicit]] for why a model can be "implicit" and still have exactly one direction.

Two things this package is for

1. It is the first thing in the project that produces a SampleBelief. LenticulumCore has declared that type since the beginning and nothing has ever constructed one. A generator is what it was for — and the moment one exists, Mycelium.combine throws, which is messages.md §1's recorded main gap made concrete.

2. A discriminator is a way around that gap. The central identity of the paper is

\[D^\ast(x) = \frac{p(x)}{p(x)+q(x)} \qquad\Longrightarrow\qquad \operatorname{logit} D^\ast(x) = \log p(x) - \log q(x)\]

so a classifier estimates a log-density ratio without either density — and a ratio is precisely what importance reweighting needs. reweight is the operation; effective_sample_size is the diagnostic that says whether it meant anything.

What it cannot do

Train adversarially. The generator descends the same quantity the discriminator ascends, and a Bethe free energy has one sign. The graph expresses the GAN's wiring exactly and its objective not at all — see [[GANs as Two Factors]] §4, which argues the missing structure is Ghani–Hedges–Winschel–Zahn's open game, a third lens-shaped object alongside the parametric lens and the statistical game.

Dependencies are LuxCore, Random, LinearAlgebra — no Lux, no AD, no training loop. This package scores and reweights; fitting the two networks is somebody else's job.

Concept notes: [[Implicit Generative Models]], [[GANs as Two Factors]], [[Three Senses of Implicit]]. Per-file notes: [[generator]], [[ratio]].

source
Adversarial.GeneratorFactor — Type
GeneratorFactor(net, dims::NamedTuple; nsamples = 64, rng = Random.default_rng(),
                channels = (:z, :x), input = identity)

$x = G_\theta(z)$ — an implicit generative model: samples come out, densities do not.

net is any LuxCore.AbstractLuxLayer mapping a latent to a sample. dims names the two channels and their dimensions. Parameters and state are the net's, untouched.

g = GeneratorFactor(my_decoder, (z = 8, x = 2); nsamples = 512)
Unidirectional, and honestly so

supported_polarities returns one element. Inverting $G_\theta$ — recovering the latent that produced a given sample — is the "GAN inversion" problem and is generally intractable; there is no residual to solve, because a generator is a function with no relation behind it. Contrast ImplicitLayers.DEQFactor, which has two.

That is not a limitation of this wrapper. It is what "implicit" means in Mohamed & Lakshminarayanan's sense as opposed to README.md's sense — see Three Senses of Implicit.md.

It emits particles, not a point

The message is a LenticulumCore.SampleBelief — the first one anything in this project has ever produced. And Mycelium.combine throws on two of those, which is messages.md §1's recorded main gap. RatioFactor is the route around it.

source
Adversarial.GeneratorModel — Type
GeneratorModel(factor)

The open model: a deterministic map with the latent as its input. Pure — a generator hides nothing given z; what it hides is the density of its own output, and that is not what $\llbracket c\rrbracket$ means.

source
Adversarial.NoiseSource — Type
NoiseSource(channel, dim; nsamples = 64, rng = Random.default_rng(), dist = randn)

The latent prior $q(z)$ as an emitting factor: draws nsamples particles and hands them on as a LenticulumCore.SampleBelief.

This is the $p(z)$ node of every GAN diagram, and it is a factor rather than a property of the variable because of [[Everything is a Factor]] — a variable is a wire and has no content of its own. Non-learnable, arity one, emitting only, exactly like Mycelium.DataFactor; the difference is that a DataFactor clamps a point and this scatters a cloud.

dist(rng, dim, n) may be replaced to sample from something other than a standard normal.

source
Adversarial.RatioFactor — Type
RatioFactor(net, dim; channel = :x, input = identity)

A density-ratio factor: a unary factor on channel whose potential is an estimated $\log r(x) = \log p(x) - \log q(x)$.

net is any LuxCore.AbstractLuxLayer mapping a sample to a logit (a real number, or a length-1 vector). It is the GAN discriminator, and it is read as a ratio estimator rather than as a classifier — the two are the same object under the identity above.

d = RatioFactor(my_critic, 2)          # a unary factor on :x
The energy is estimated, not computed

$\mathbf{l}^c(x) = -\log r(x)$, so $\mathbb{E}_{q}[-\log r] = \mathrm{KL}(q\,\|\,p)$ — the factor's energy is the integrand of a KL divergence. But $\log r$ comes from a fitted network, so unlike every other factor in this project the energy itself is an estimate.

Bayesian Lens.md licenses inexact inversions. This is a different thing: an inexact energy, and the free energy has no term that accounts for it. See ratio.md §5.

source
Adversarial.RatioModel — Type
RatioModel(factor)

The open model of a unary evidence factor. There is no forward map — a ratio factor does not produce an x, it scores one — so the inversion is where all the content is.

source
Adversarial.discriminator — Method
discriminator(f::RatioFactor, x, ps, st) -> (Real, st)

$D(x) = \sigma(\log r(x)) = p(x)/(p(x)+q(x))$ — the classifier reading of the same number.

Provided so the identity is visible in code: discriminator and logratio are the logistic transform of one another, and which one you call is a matter of which paper you are reading.

source
Adversarial.effective_sample_size — Method
effective_sample_size(b::SampleBelief) -> Real

$\mathrm{ESS} = 1/\sum_i w_i^2$ for normalised weights; length(samples) when unweighted.

Ranges from 1 (one particle carries everything — the answer is a point estimate wearing a distribution's clothes) to n (uniform). Report it or the reweighting means nothing, and it is the honest reason a ratio-based combine is not simply dropped into Mycelium.combine: the operation is defined but its quality is not, and a combine that silently degenerates is worse than one that throws.

source
Adversarial.generate — Method
generate(f::GeneratorFactor, z, ps, st) -> (x, st)

One forward pass: $G_\theta(z)$ for a single latent.

source
Adversarial.logratio — Method
logratio(f::RatioFactor, x, ps, st) -> (Real, st)

$\log r(x) = \log p(x) - \log q(x)$, i.e. the discriminator's logit.

source
Adversarial.reweight — Method
reweight(f::RatioFactor, b::SampleBelief, ps, st) -> (SampleBelief, st)

Importance-reweight a particle set by the estimated ratio: $w_i \propto w_i^{\text{old}}\, r(x_i)$.

If b's particles are drawn from $q$ and $r = p/q$, the result is a weighted particle approximation of $p$. This is the operation messages.md §1 says is missing — pooling a sample-based belief with another distribution — performed by the one object in the graph that knows the ratio.

Weights are normalised and computed in log space (a shift by the maximum before exponentiating), because raw ratios overflow for any interesting p/q pair.

[!warning] Reweighting concentrates The variance of the weights grows with the divergence between p and q, and the particle set silently collapses onto a handful of samples. Check effective_sample_size; it is the diagnostic that says whether the answer means anything.

source
Adversarial.weighted_mean — Method
weighted_mean(b::SampleBelief)
weighted_cov(b::SampleBelief)

Self-normalised importance-sampling estimates of the first two moments.

These are the only way to read anything out of a SampleBelief, and they are how the test suite checks reweighting against a Gaussian closed form.

source
LenticulumCore.invert — Method
LenticulumCore.invert(lens, π, inputs, ps, st)

Reweight the prior π by the estimated ratio.

AmortisedInversion is the right label: the inversion is a learned network with its own parameters, which lens.md says is the structural reason a factor cannot be a Lux layer. A RatioFactor is the first factor in the project where that description is literally true — the discriminator is trained separately from whatever produced the samples.

[!warning] The prior is consumed, so this is a posterior A unary factor has no other channel to read, so the only thing to reweight is the incoming belief — which makes the message a posterior rather than a likelihood, and double-counts on a variable of degree > 1. The same wall the diffusion and DEQ factors hit, for the same reason: there is nothing to divide out of a neural network. See ratio.md §6.

source
LenticulumCore.pushforward — Method
LenticulumCore.pushforward(f::GeneratorFactor, belief, ps, st) -> (belief, st)
LenticulumCore.pushforward(m::GeneratorModel, belief, ps, st) -> (belief, st)

Push a belief on the latent channel through $G_\theta$ — AutoBayes' $c_*\pi$.

[!note] This is the first implementation of pushforward in the project LenticulumCore has declared forward, logdensity and pushforward since the beginning and nothing has implemented the last one. A generator is the case where it is easy: the map is deterministic, so pushing a particle set forward is mapping over it, with no marginalisation integral at all. open_model.md warns that pushforward is "one of the two expensive operations"; for a sampler it is the cheap one, which is the whole appeal of implicit generative models.

inoutwhy
DiracBeliefDiracBeliefa function of a point is a point
SampleBeliefSampleBeliefmap each particle; weights are carried through unchanged
TrivialBeliefTrivialBeliefno latent, no samples

The middle row is the interesting one and the weight-preservation is not a detail: pushing a weighted particle set through a deterministic map leaves the weights alone, because the map is a bijection on particle indices. That is why importance weights survive a generator and why RatioFactor can be applied downstream of one.

source
Mycelium.factor_message — Method
Mycelium.factor_message(f::RatioFactor, target, polarity, inputs, prior, ps, st)

The reweighted belief on f's single channel. Falls back to the prior when no message has arrived — which for a unary factor is the normal case, since its only neighbour is its target.

source
Mycelium.factor_message — Method
Mycelium.factor_message(f::GeneratorFactor, target, polarity, inputs, prior, ps, st)

The pushforward of the latent belief. Asking for a message on the latent channel throws: that is GAN inversion, and this factor cannot do it.

source
Mycelium.local_free_energy — Method
Mycelium.local_free_energy(f::RatioFactor, msgs, ps, st)

$\mathbb{E}_{b}[-\log r(x)]$ under the incoming particle set — a Monte-Carlo estimate of $\mathrm{KL}(q\,\|\,p)$ when r is well fitted.

This is the number a GAN's generator descends, and the sign is worth dwelling on: the discriminator ascends the same quantity. A Bethe free energy has one sign, so the graph can express half of a GAN. See GANs as Two Factors.md §4.

source