Missing ties#
Network data often have dyads whose value is unknown: survey respondents who
didn’t answer, people who left before the second wave, ties that couldn’t be
coded. Treating them as absent biases the density, and the rest of the
model, downwards. ergmx treats them as R’s ergm does.
Marking missing dyads#
A missing dyad is an edge with a true na attribute, as in R’s network
package (net[i, j] <- NA). In igraph:
import ergmx
from ergmx import datasets
samplk3 = datasets.load("samplk3")
# Two monks didn't answer: their nominations are unknown.
nonrespondents = [0, 6]
g = samplk3.copy()
g.delete_edges(g.es.select(_source_in=nonrespondents))
g.es["na"] = False
g.add_edges([(i, j) for i in nonrespondents for j in range(g.vcount()) if j != i],
attributes={"na": True})
print(g.ecount() - sum(g.es["na"]), "ties,", sum(g.es["na"]), "missing dyads")
50 ties, 34 missing dyads
In networkx: G.add_edge(u, v, na=True). Statistics (summary_stats) count
missing dyads as non-ties, as R’s summary() does.
Fitting#
fit = ergmx.ergm(g, "edges + mutual", seed=1)
fit.summary()
Monte Carlo Maximum Likelihood Results:
Estimate Std. Error MCMC % z value Pr(>|z|)
edges -1.9538 0.2257 0 -8.655 <1e-04 ***
mutual 1.8006 0.5377 0 3.349 0.00081 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Log-likelihood: -124.5439 (MC SE 0.050) AIC: 253.0879 BIC: 260.2995
Missing dyads: 34, assumed missing at random.
Converged after 3 iterations (4 chains, 1024 samples).
R’s ergm gives −1.98 and 1.82 for this model, within Monte Carlo error. With the complete network, the estimates are −2.15 and 2.30, with smaller standard errors (0.22 and 0.48): the 34 unknown dyads carry no information, so the two monks’ nominations, and how many of them were returned, are inferred from the model.
The method#
The likelihood is that of the observed dyads, summing over every possible value of the missing ones (Handcock and Gile 2010):
It assumes the dyads are missing at random: whether a dyad is missing doesn’t depend on its unobserved value. A student who skipped the survey because they had few friends breaks the assumption, and no method can tell from the data alone.
Estimating equation. The MLE solves \(E_\theta[g(Y) \mid y_{obs}] = E_\theta[g(Y)]\): the expected statistics given the observed dyads equal the expected statistics overall. Each Monte Carlo iteration draws two samples, one unconditional and one in which only the missing dyads change, and takes the Newton step \((\Sigma - \Sigma_{obs})\,\delta = \mu_{obs} - \mu\).
Standard errors come from the information \(\Sigma - \Sigma_{obs}\) (the missing information principle, Louis 1982), plus the MCMC error of both samples.
Dyad-independent models are still fitted exactly: by logistic regression on the observed dyads, as the missing ones factor out.
Starting values are the MPLE on the observed dyads, with missing dyads as non-ties in the change statistics. ergm imputes them at random first, so its MPLE differs a little between runs; the MLE is the same.
The log-likelihood is that of the observed dyads, by path sampling with conditional samples along the path, so AIC, BIC and
ergmx.compare()work as usual.
Goodness of fit#
With missing dyads, ergmx.gof() compares the simulated networks with
networks imputed from the model given the observed dyads, averaged, rather
than with the network with missing dyads as non-ties, which would make every
model look like it predicts too many ties:
fit.gof(seed=1)["odegree"]
Goodness-of-fit for out-degree
obs min mean max MC p-value
0 0.04 0 0.52 3 0.86
1 0.05 0 2.19 7 0.34
2 0.31 1 3.82 9 0.00
3 14.55 0 4.39 11 0.00
4 2.56 0 3.64 9 0.52
5 0.22 0 2.08 6 0.26
6 0.14 0 0.97 4 0.78
7 0.09 0 0.29 2 0.50
8 0.03 0 0.09 1 0.18
9 0.01 0 0.01 1 0.02
The observed values are averages over the imputations, so they need not be
whole numbers. They show what edges + mutual misses: nearly every monk named
three brothers, the fixed-choice design that the constraint bd(maxout=4),
or odegrees, represents (see Constraints).
Several networks, and waves#
In networks combined with ergmx.Networks(), each network’s missing
dyads are missing in the joint model, and ergmx.gofN() imputes them
too. In a series of networks (ergmx.tergm()), the missing dyads of the
networks transitioned to are missing as here; those of the networks
transitioned from, which the next transition is conditional on, are
imputed first with na_impute=, as tergm’s NA.impute: see
Missing dyads.