Constraints#
A model’s sample space is the set of networks it puts probability on: by default every network on the vertices. Constraints restrict it, for when some networks could not have been observed, or to condition on some features of the network. They are given as in R’s ergm:
ergmx.ergm(network, formula, constraints="bd(maxout=4)")
ergmx.ergm(network, formula, constraints="~degrees") # a leading ~ is fine
ergmx.ergm(network, formula, constraints="bd(maxout=4) + blocks('level', levels2=2)")
Constraint |
Networks allowed |
|---|---|
|
Degrees within bounds (one value, or one per vertex). Undirected networks use |
|
The dyads of some mixing types of |
|
Bounds by the alters’ classes: |
|
Ties only between vertices with the same value of |
|
With |
|
These dyads are fixed ( |
|
Every dyad but these is fixed. |
|
The observed dyads are fixed: only the missing ones vary, to simulate them. |
|
Every vertex keeps its degree (its in- and out-degrees, if directed). |
|
Every vertex keeps its out-degree, or its in-degree (directed networks). |
|
The vertices of the first (second) mode of a bipartite network keep their degrees. |
|
The number of edges is kept. |
Dyads are given as an edge list of R’s vertex numbers, from 1
(matrix(c(1, 9, 2, 6), ncol=2, byrow=TRUE), or a list of pairs), a
logical \(n \times n\) matrix, a graph, or the name of a graph attribute
holding one. R’s argument names with dots (fixed.dyads) work in strings;
in Python they have underscores.
Bounded degrees: fixed-choice designs#
In Sampson’s monastery, each monk named the three brothers he liked most (four when he couldn’t choose). No network with more than four nominations per monk could have been observed, so the model shouldn’t put probability on any:
import ergmx
from ergmx import datasets
samplk3 = datasets.load("samplk3")
bounded = ergmx.ergm(samplk3, "edges + mutual", constraints="bd(maxout=4)", seed=1)
bounded.summary()
Monte Carlo Maximum Likelihood Results:
Estimate Std. Error MCMC % z value Pr(>|z|)
edges -1.6356 0.2924 0 -5.593 <1e-04 ***
mutual 2.2863 0.4836 0 4.728 <1e-04 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Log-likelihood, relative to the null model: 17.6318 (MC SE 0.070) AIC: -31.2635 BIC: -23.8164
Constraints: bd(maxout=4).
Converged after 3 iterations (4 chains, 1024 samples).
Without the bound, the model explains the low density with a large negative
edges coefficient; with it, much of the sparsity comes from the design,
and edges is closer to zero (compare R’s ergm: −1.64 and 2.32).
Fixing blocks of dyads#
blocks fixes the dyads of chosen mixing types. With levels2=-1, every
mixing type but the first is fixed, so only the ties among 7th graders are
modeled; the dyads between them and everyone else keep their observed
(absent) values:
mesa = datasets.load("faux.mesa.high")
fit = ergmx.ergm(mesa, "edges + nodematch('Race')", constraints="blocks('Grade', levels2=-1)")
fit.summary()
Maximum Likelihood Results:
Estimate Std. Error MCMC % z value Pr(>|z|)
edges -3.2442 0.1632 0 -19.876 <1e-04 ***
nodematch.Race 0.1233 0.2359 0 0.523 0.60119
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Log-likelihood: -315.4093 AIC: 634.8186 BIC: 645.9083
Constraints: blocks('Grade', levels2=-1).
This model is dyad-independent and blocks fixes dyads independently of
each other, so the fit is still an exact logistic regression, over the 1,891
pairs of 7th graders.
Preserving degrees#
Conditioning on the degrees asks whether a network has more triangles, or more ties within groups, than networks with the same degrees: degree heterogeneity alone can produce clustering.
fit = ergmx.ergm(mesa, "nodematch('Grade') + gwesp(0.5, fixed=TRUE)", constraints="degrees",
seed=1)
fit.summary()
Monte Carlo Maximum Likelihood Results:
Estimate Std. Error MCMC % z value Pr(>|z|)
nodematch.Grade 2.1724 0.1676 0 12.964 <1e-04 ***
gwesp.fixed.0.5 1.3756 0.1279 0 10.754 <1e-04 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Log-likelihood, relative to the null model: 252.2093 (MC SE 0.556) AIC: -500.4187 BIC: -484.5227
Constraints: degrees.
Converged after 3 iterations (4 chains, 1024 samples).
Same-grade ties and triadic closure remain strong given the degrees. Terms
whose statistics the constraint keeps constant, such as edges, kstar,
gwdegree or nodefactor under degrees, can’t be estimated: ergmx
warns, fixes their coefficients at 0 and reports them as constant.
The MCMC uses moves that keep the degrees: it swaps the endpoints of two
ties and, in directed networks, also reverses cyclic triples (\(i \to j \to k
\to i\)), without which some networks with the same degrees can’t reach each
other. Under b1degrees (b2degrees) it moves the other end of a tie to
another vertex of that mode, and under edges it swaps a tie for a non-tie.
What changes with dyad-dependent constraints#
bd, edges and the degree constraints are dyad-dependent: whether a
dyad may change depends on the others. Then:
the MPLE ignores the constraint, so
estimate="MPLE"is not available, and the Monte Carlo MLE starts from the contrastive divergence estimate, which respects it (init="CD"), as in ergm;dyad-independent models still need the Monte Carlo MLE;
the log-likelihood is relative to the null model, the uniform distribution over the allowed networks, as in ergm. The summary says so. Relative log-likelihoods, AICs and BICs compare models with the same constraints, and
ergmx.compare()refuses models with different ones.
ergmx.simulate() and ergmx.gof() take the same constraints, and
a fit’s simulations and goodness of fit use its own.