Missing ties#

Network data often have dyads whose value is unknown: survey respondents who didn’t answer, people who left before the second wave, ties that couldn’t be coded. Treating them as absent biases the density, and the rest of the model, downwards. ergmx treats them as R’s ergm does.

Marking missing dyads#

A missing dyad is an edge with a true na attribute, as in R’s network package (net[i, j] <- NA). In igraph:

import ergmx
from ergmx import datasets

samplk3 = datasets.load("samplk3")
# Two monks didn't answer: their nominations are unknown.
nonrespondents = [0, 6]
g = samplk3.copy()
g.delete_edges(g.es.select(_source_in=nonrespondents))
g.es["na"] = False
g.add_edges([(i, j) for i in nonrespondents for j in range(g.vcount()) if j != i],
            attributes={"na": True})
print(g.ecount() - sum(g.es["na"]), "ties,", sum(g.es["na"]), "missing dyads")
50 ties, 34 missing dyads

In networkx: G.add_edge(u, v, na=True). Statistics (summary_stats) count missing dyads as non-ties, as R’s summary() does.

Fitting#

fit = ergmx.ergm(g, "edges + mutual", seed=1)
fit.summary()
Monte Carlo Maximum Likelihood Results:

         Estimate  Std. Error  MCMC %  z value  Pr(>|z|)
edges     -1.9538      0.2257       0   -8.655    <1e-04 ***
mutual     1.8006      0.5377       0    3.349   0.00081 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Log-likelihood: -124.5439 (MC SE 0.050)   AIC: 253.0879   BIC: 260.2995
Missing dyads: 34, assumed missing at random.
Converged after 3 iterations (4 chains, 1024 samples).

R’s ergm gives −1.98 and 1.82 for this model, within Monte Carlo error. With the complete network, the estimates are −2.15 and 2.30, with smaller standard errors (0.22 and 0.48): the 34 unknown dyads carry no information, so the two monks’ nominations, and how many of them were returned, are inferred from the model.

The method#

The likelihood is that of the observed dyads, summing over every possible value of the missing ones (Handcock and Gile 2010):

\[ L(\theta) = \sum_{y_{mis}} P_\theta(y_{obs}, y_{mis}). \]

It assumes the dyads are missing at random: whether a dyad is missing doesn’t depend on its unobserved value. A student who skipped the survey because they had few friends breaks the assumption, and no method can tell from the data alone.

  • Estimating equation. The MLE solves \(E_\theta[g(Y) \mid y_{obs}] = E_\theta[g(Y)]\): the expected statistics given the observed dyads equal the expected statistics overall. Each Monte Carlo iteration draws two samples, one unconditional and one in which only the missing dyads change, and takes the Newton step \((\Sigma - \Sigma_{obs})\,\delta = \mu_{obs} - \mu\).

  • Standard errors come from the information \(\Sigma - \Sigma_{obs}\) (the missing information principle, Louis 1982), plus the MCMC error of both samples.

  • Dyad-independent models are still fitted exactly: by logistic regression on the observed dyads, as the missing ones factor out.

  • Starting values are the MPLE on the observed dyads, with missing dyads as non-ties in the change statistics. ergm imputes them at random first, so its MPLE differs a little between runs; the MLE is the same.

  • The log-likelihood is that of the observed dyads, by path sampling with conditional samples along the path, so AIC, BIC and ergmx.compare() work as usual.

Goodness of fit#

With missing dyads, ergmx.gof() compares the simulated networks with networks imputed from the model given the observed dyads, averaged, rather than with the network with missing dyads as non-ties, which would make every model look like it predicts too many ties:

fit.gof(seed=1)["odegree"]
Goodness-of-fit for out-degree

        obs       min       mean       max  MC p-value
0      0.04         0       0.52         3        0.86
1      0.05         0       2.19         7        0.34
2      0.31         1       3.82         9        0.00
3     14.55         0       4.39        11        0.00
4      2.56         0       3.64         9        0.52
5      0.22         0       2.08         6        0.26
6      0.14         0       0.97         4        0.78
7      0.09         0       0.29         2        0.50
8      0.03         0       0.09         1        0.18
9      0.01         0       0.01         1        0.02

The observed values are averages over the imputations, so they need not be whole numbers. They show what edges + mutual misses: nearly every monk named three brothers, the fixed-choice design that the constraint bd(maxout=4), or odegrees, represents (see Constraints).

Several networks, and waves#

In networks combined with ergmx.Networks(), each network’s missing dyads are missing in the joint model, and ergmx.gofN() imputes them too. In a series of networks (ergmx.tergm()), the missing dyads of the networks transitioned to are missing as here; those of the networks transitioned from, which the next transition is conditional on, are imputed first with na_impute=, as tergm’s NA.impute: see Missing dyads.