An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.
Running a network study
A network study succeeds or fails at the design stage. Four decisions do most of the work. The question, which determines everything after it. The boundary, meaning exactly who is inside the network and who is not. The relation, which has to be one specific tie definition rather than a vague sense of connection. And the collection method, whether a roster, a name generator or digital trace data. Missing actors do far more damage here than missing rows do in ordinary survey work, because every tie incident to the missing actor goes with them, along with every path that ran through them.
Why it matters
Ordinary survey work loses a respondent and loses one row. A network study loses a respondent and loses the ties incident to that person, plus every path that ran through them, so two people you did collect can look further apart than they are. That is why design comes first. You cannot patch a network dataset the way you patch a spreadsheet, because what is missing is not the cell, it is the connection.
An analyst achieves a 70 per cent response rate on a 60-person network survey, records a tie only when both actors name each other, and reports betweenness scores as final. What is the strongest objection?
Formulas
Worked examples
A hospital asks for a study of the network behind handover quality. The obvious boundary is the ward roster of 46 nurses.
The roster is a nominalist boundary, drawn by the researcher from an administrative list. It is defensible and it will miss the two agency staff who cover most night shifts and the pharmacist everyone phones. A realist alternative asks participants who they actually rely on and includes whoever is named often enough. Each approach produces a different network and a different answer about fragility. Neither is wrong. What is wrong is failing to state which one was used, because the reader cannot otherwise judge the density or the bridges.
A firm has three years of internal messaging logs and wants a collaboration network built from them without running a survey.
Trace data is complete, cheap and measures something different from what people report. A message log records contact rather than reliance, over-counts scheduling chatter, and misses conversations held in person. It also leaves a consent question that a survey answers explicitly. A workable design builds the network from the logs, validates it against a short survey on a sample of staff, reports the agreement rate, and defines a tie by a stated threshold such as a minimum of two-way exchanges per month.
Common mistakes
- ✗A network study is a survey with extra questions. The dependency structure changes the statistics. Standard significance tests assume independent observations, and network data violates that by construction, which is why permutation approaches such as the quadratic assignment procedure are used instead of ordinary tests.
- ✗Removing names protects participants. Removing names does not remove position. Someone who is the only bridge between two departments is identifiable from the structure alone, and small networks are effectively re-identifiable from the pattern of degrees. Protection has to come from aggregation, access control and informed consent, not from relabelling.
- ✗A high response rate makes the network safe. Coverage of pairs falls faster than coverage of people. Under mutual confirmation, even at 90 per cent response roughly a fifth of pairs have a missing end, and the actors most likely to be missing, such as contractors and part-timers, are often the ones holding the bridges.
- ✗Digital trace data removes the measurement problem. Trace data measures the channel it was logged on and nothing else. It records contact rather than reliance, it stops at the platform boundary, and it silently excludes anyone who works mostly offline. It is a different instrument, not a neutral one.
Revision bullets
- •Design order: question, then boundary, then relation, then collection method
- •Nominalist boundaries come from a list; realist boundaries come from the actors
- •One relation per network; never merge advice, friendship and reporting lines
- •Under mutual confirmation, dyad coverage is roughly the square of the response rate
- •Structural position re-identifies people even after names are removed
- •Trace data and self-report measure different relations, so validate one against the other
Quick check
An analyst achieves a 70 per cent response rate on a 60-person network survey, records a tie only when both actors name each other, and reports betweenness scores as final. What is the strongest objection?
A study of a 30-person team publishes its network diagram with names replaced by codes. A team member reads the report. What is the realistic privacy risk?
Connected topics
More in Networks
Sources
- Laumann, Marsden & Prensky (1983)Laumann, E. O., Marsden, P. V., & Prensky, D. "The Boundary Specification Problem in Network Analysis." In R. S. Burt & M. J. Minor (eds), Applied Network Analysis: A Methodological Introduction, SAGE, 1983, pp. 18-34.The standard statement of the nominalist and realist approaches to drawing a network boundary.
- Marsden, P. V. "Network Data and Measurement." Annual Review of Sociology, 16, 435-463, 1990.Reviews name generators, rosters and informant accuracy, and what each instrument actually measures.
- Kossinets, G. "Effects of Missing Data in Social Networks." Social Networks, 28(3), 247-268, 2006.Quantifies how non-response and boundary error distort network measures.
- Kadushin, C. "Who Benefits from Network Analysis: Ethics of Social Network Research." Social Networks, 27(2), 139-153, 2005.Sets out why consent and anonymity work differently when respondents name other people.