Constraints#

A model’s sample space is the set of networks it puts probability on: by default every network on the vertices. Constraints restrict it, for when some networks could not have been observed, or to condition on some features of the network. They are given as in R’s ergm:

ergmx.ergm(network, formula, constraints="bd(maxout=4)")
ergmx.ergm(network, formula, constraints="~degrees")              # a leading ~ is fine
ergmx.ergm(network, formula, constraints="bd(maxout=4) + blocks('level', levels2=2)")

Constraint

Networks allowed

bd(maxout=, maxin=, minout=, minin=)

Degrees within bounds (one value, or one per vertex). Undirected networks use minout and maxout for degrees.

blocks(attr, levels=, levels2=)

The dyads of some mixing types of attr are fixed at their observed values: those nodemix(attr, levels, levels2) would count. levels2 selects them as for nodemix; by default none.

bd(attribs=, maxout=, ...)

Bounds by the alters’ classes: attribs (a graph attribute, or a vertices x classes logical matrix) marks each vertex’s classes, and the bounds (same shape) limit each vertex’s ties to alters of each class.

blockdiag(attr)

Ties only between vertices with the same value of attr: the dyads between blocks are fixed at no tie.

Dyads(fix=~terms, vary=~terms)

With fix, the dyads that the dyad-independent terms count are fixed; with vary, only those may vary; with both, the dyads either lets vary.

fixedas(fixed.dyads=, present=, absent=)

These dyads are fixed (present checked to be ties, absent non-ties).

fixallbut(free.dyads)

Every dyad but these is fixed.

observed

The observed dyads are fixed: only the missing ones vary, to simulate them.

degrees

Every vertex keeps its degree (its in- and out-degrees, if directed).

odegrees, idegrees

Every vertex keeps its out-degree, or its in-degree (directed networks).

b1degrees, b2degrees

The vertices of the first (second) mode of a bipartite network keep their degrees.

edges

The number of edges is kept.

Dyads are given as an edge list of R’s vertex numbers, from 1 (matrix(c(1, 9, 2, 6), ncol=2, byrow=TRUE), or a list of pairs), a logical \(n \times n\) matrix, a graph, or the name of a graph attribute holding one. R’s argument names with dots (fixed.dyads) work in strings; in Python they have underscores.

Bounded degrees: fixed-choice designs#

In Sampson’s monastery, each monk named the three brothers he liked most (four when he couldn’t choose). No network with more than four nominations per monk could have been observed, so the model shouldn’t put probability on any:

import ergmx
from ergmx import datasets

samplk3 = datasets.load("samplk3")
bounded = ergmx.ergm(samplk3, "edges + mutual", constraints="bd(maxout=4)", seed=1)
bounded.summary()
Monte Carlo Maximum Likelihood Results:

         Estimate  Std. Error  MCMC %  z value  Pr(>|z|)
edges     -1.6356      0.2924       0   -5.593    <1e-04 ***
mutual     2.2863      0.4836       0    4.728    <1e-04 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Log-likelihood, relative to the null model: 17.6318 (MC SE 0.070)   AIC: -31.2635   BIC: -23.8164
Constraints: bd(maxout=4).
Converged after 3 iterations (4 chains, 1024 samples).

Without the bound, the model explains the low density with a large negative edges coefficient; with it, much of the sparsity comes from the design, and edges is closer to zero (compare R’s ergm: −1.64 and 2.32).

Fixing blocks of dyads#

blocks fixes the dyads of chosen mixing types. With levels2=-1, every mixing type but the first is fixed, so only the ties among 7th graders are modeled; the dyads between them and everyone else keep their observed (absent) values:

mesa = datasets.load("faux.mesa.high")
fit = ergmx.ergm(mesa, "edges + nodematch('Race')", constraints="blocks('Grade', levels2=-1)")
fit.summary()
Maximum Likelihood Results:

                 Estimate  Std. Error  MCMC %  z value  Pr(>|z|)
edges             -3.2442      0.1632       0  -19.876    <1e-04 ***
nodematch.Race     0.1233      0.2359       0    0.523   0.60119 
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Log-likelihood: -315.4093   AIC: 634.8186   BIC: 645.9083
Constraints: blocks('Grade', levels2=-1).

This model is dyad-independent and blocks fixes dyads independently of each other, so the fit is still an exact logistic regression, over the 1,891 pairs of 7th graders.

Preserving degrees#

Conditioning on the degrees asks whether a network has more triangles, or more ties within groups, than networks with the same degrees: degree heterogeneity alone can produce clustering.

fit = ergmx.ergm(mesa, "nodematch('Grade') + gwesp(0.5, fixed=TRUE)", constraints="degrees",
                 seed=1)
fit.summary()
Monte Carlo Maximum Likelihood Results:

                  Estimate  Std. Error  MCMC %  z value  Pr(>|z|)
nodematch.Grade     2.1724      0.1676       0   12.964    <1e-04 ***
gwesp.fixed.0.5     1.3756      0.1279       0   10.754    <1e-04 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Log-likelihood, relative to the null model: 252.2093 (MC SE 0.556)   AIC: -500.4187   BIC: -484.5227
Constraints: degrees.
Converged after 3 iterations (4 chains, 1024 samples).

Same-grade ties and triadic closure remain strong given the degrees. Terms whose statistics the constraint keeps constant, such as edges, kstar, gwdegree or nodefactor under degrees, can’t be estimated: ergmx warns, fixes their coefficients at 0 and reports them as constant.

The MCMC uses moves that keep the degrees: it swaps the endpoints of two ties and, in directed networks, also reverses cyclic triples (\(i \to j \to k \to i\)), without which some networks with the same degrees can’t reach each other. Under b1degrees (b2degrees) it moves the other end of a tie to another vertex of that mode, and under edges it swaps a tie for a non-tie.

What changes with dyad-dependent constraints#

bd, edges and the degree constraints are dyad-dependent: whether a dyad may change depends on the others. Then:

  • the MPLE ignores the constraint, so estimate="MPLE" is not available, and the Monte Carlo MLE starts from the contrastive divergence estimate, which respects it (init="CD"), as in ergm;

  • dyad-independent models still need the Monte Carlo MLE;

  • the log-likelihood is relative to the null model, the uniform distribution over the allowed networks, as in ergm. The summary says so. Relative log-likelihoods, AICs and BICs compare models with the same constraints, and ergmx.compare() refuses models with different ones.

ergmx.simulate() and ergmx.gof() take the same constraints, and a fit’s simulations and goodness of fit use its own.