MEIRU Microdata Portal
  • Home
  • Overview
  • Microdata Catalog
  • Citations
  • Login
    Login
    Home / Central Data Catalog / MWI-MEIRU-SPOUSELINKS-2002-V01
central

Cleaned spouse and marriage data

Malawi
Get Microdata
Reference ID
MWI-MEIRU-SPOUSELINKS-2002-v01
Producer(s)
Professor Amelia (Mia) Crampin
Collections
Karonga HDSS Combined / Harmonized Datasets
Metadata
Documentation in PDF DDI/XML JSON
Created on
Nov 14, 2025
Last modified
Nov 19, 2025
Page views
6687
Downloads
106
  • Study Description
  • Data Dictionary
  • Downloads
  • Get Microdata
  • Identification
  • Version
  • Producers and sponsors
  • Data collection
  • Access policy
  • Data Access
  • Metadata production
  • Identification

    Survey ID number

    MWI-MEIRU-SPOUSELINKS-2002-v01

    Title

    Cleaned spouse and marriage data

    Country
    Name Country code
    Malawi MWI
    Study type

    Sample Survey

    Abstract

    The do-file marital_spouselinks.do combines all data on people's marital statuses and reported spouses to create the following datasets:

    1. all_marital_reports - a listing of all the times an individual has reported their current marital status with the id numbers of the reported spouse(s); this listing is as reported so may include discrepancies (i.e. a 'Never married' status following a 'Married' one)
    2. all_spouse_pairs_full - a listing of each time each spouse pair has been reported plus summary information on co-residency for each pair
    3. all_spouse_pairs_clean_summarised - this summarises the data from all_spouse_pairs_full to give start and end dates of unions
    4. marital_status_episodes - this combines data from all the sources to create episodes of marital status, each has a start and end date and a marital status, and if currently married, the spouse ids of the current spouse(s) if reported. There are several variables to indicate where each piece of information is coming from.

    The first 2 datasets are made available in case people need the 'raw' data for any reason (i.e. if they only want data from one study) or if they wish to summarise the data in a different way to what is done for the last 2 datasets.

    The do-file is quite complicated with many sources of data going through multiple processes to create variables in the datasets so it is not always straightforward to explain where each variable come from on the documentation. The 4 datasets build on each other and the do-file is documented throughout so anyone wanting to understand in great detail may be better off examining that. However, below is a brief description of how the datasets are created:

    Marital status data are stored in the tables of the study they were collected in:
    AHS Adult Health Study [ahs_ahs1]
    CEN Census (initial CRS census) [cen_individ]
    CENM In-migration (CRS migration form) [crs_cenm]
    GP General form (filled for various reasons) [gp_gpform]
    SEI Socio-economic individual (annual survey from 2007 onwards) [css_sei]
    TBH TB household (study of household contacts of TB patients) [tb_tbh]
    TBO TB controls (matched controls for TB patients) [tb_tbo & tb_tboto2007]
    TBX TB cases (TB patients) [tb_tbx & tb_tbxto2007]
    In many of the above surveys as well as their current marital status, people were asked to report their current and past spouses along with (sometimes) some information about the marriage (start/end year etc.). These data are stored all together on the table gen_spouse, with variables indicating which study the data came from.
    Further evidence of spousal relationships is taken from gen_identity (if a couple appear as co-parents to a CRS member) and from crs_residency_episodes_clean_poly, a combined dataset (if they are living in the same household at the same time). Note that co-parent couples who are not reported in gen_spouse are only retained in the datasets if they have co-resident episodes.

    The marital status data are appended together and the spouse id data merged in. Minimal data editing/cleaning is carried out. As the spouse data are in long format, this dataset is reshaped wide to have one line per marital status report (polygamy in the area allows for men to have multiple spouses at one time): this dataset is saved as all_marital_reports.

    The list of reported spouses on gen_spouse is appended to a list of co-parents (from gen_identity) and this list is cleaned to try to identify and remove obvious id errors (incestuous links, same sex [these are not reported in this culture] and large age difference). Data reported by men and women are compared and variables created to show whether one or both of the couple report the union.
    Many records have information on start and end year of marriage, and all have the date the union was reported. This listing is compared to data from residency episodes to add dates that couples were living together (not all have start/end dates so this is to try to supplement this), in addition the dates that each member of the couple was last known to be alive or first known to be dead are added (from the residency data as well). This dataset with all the records available for each spouse pair is saved as all_spouse_pairs_full.

    The date data from all_spouse_pairs_full are then summarised to get one line per couple with earliest and latest known married date for all, and, if available, marriage and separation date. For each date there are also variables created to indicate the source of the data.
    As culture only allows for women having one spouse at a time, records for women with 'overlapping' husbands are cleaned. This dataset is then saved as all_spouse_pairs_clean_summarised.

    Both the cleaned spouse pairs and the cleaned marital status datasets are converted into episodes: the spouse listing using the marriage or first known married date as the beginning and the last known married plus a year or separation date as the end, the marital status data records collapsed into periods of the same status being reported (following some cleaning to remove impossible reports) and the start date being the first of these reports, the end date being the last of the reports plus a year. These episodes are appended together and a series of processes run several times to remove overalapping episodes. To be able to assign specific spouse ids to each married episode, some episodes need to be 'split' into more than one (i.e. if a man is married to one woman from 2005 to 2017 and then marries another woman in 2008 and remains married to her till 2017 his intial married episode would be from 2005 to 2017, but this would need to be split into one from 2005 to 2008 which would just have 1 idspouse attached and another from 2008 to 2017, which would have 2 idspouse attached). After this splitting process the spouse ids are merged in.
    The final episode dataset is saved as marital_status_episodes.

    Unit of Analysis

    Individual

    Version

    Version Description

    v1: Edited data, first version, for internal use only

    Version Date

    2020-07-23

    Producers and sponsors

    Primary investigators
    Name Affiliation
    Professor Amelia (Mia) Crampin MEIRU, London School of Hygiene and Tropical Medicine (LSHTM)

    Data collection

    Mode of data collection
    • Face-to-face [f2f]
    Data Collectors
    Name Abbreviation
    Malawi Epidemiology and Intervention Research Unit MEIRU

    Access policy

    Archive where study is originally stored

    MEIRU

    Data Access

    Access authority
    Name
    Malawi Epidemiology and Intervention Research Unit
    Access conditions

    This data is made available for licensed access under the following conditions:

    1. Data and other material provided by MEIRU will not be redistributed or sold to other individuals, institutions or organisations without MEIRU's written agreement.

    2. In the case of multi-centre datasets, data originating from a single contributing member centre of the collaboration may not be analysed or reported on in isolation without the express permission of the member centre concerned.

    3. No attempt will be made to re-identify respondents, and there will be no use of the identity of any person or establishment discovered inadvertently. Any such discovery will be reported immediately to MEIRU.

    4. No attempt will be made to produce links between datasets provided by MEIRU or between MEIRU data and other datasets that could identify individuals.

    5. Any books, articles, conference papers, theses, dissertations, reports or other publications employing data obtained from MEIRU will cite the source, in line with the citation requirement provided with the dataset.

    6. An electronic copy of all publications based on the requested data will be sent to MEIRU.

    7. MEIRU, MEIRU research collaborators and the relevant funding agencies bear no responsibility for the data's use or interpretation or inferences based upon it.

    Metadata production

    DDI Document ID

    DDI-MWI-MEIRU-SPOUSELINKS-2002-v01

    Producers
    Name Abbreviation Affiliation Role
    Malawi Epidemiology and Intervention Research Unit MEIRU Agency
    Estelle McLean EM LSHTM and MEIRU Production of the documentation used to create this DDI document.
    Chifundo Kanjala CK MEIRU Data documentalist
    Jacky Saul JS LSHTM and MEIRU Production of Data dictionaries
    Kieth Branson KB LSHTM and MEIRU Production of Data dictionaries
    Joseph Kamanga JK MEIRU Metadata entry
    Dominic Nzundah DN MEIRU Metadata entry and editing
    Date of Metadata Production

    2020-07-23

    Metadata version

    DDI Document version

    version 1 (July, 2020)

    Back to Catalog
    MEIRU Microdata Portal

    © MEIRU Microdata Portal | admin@meirudata.mw | All Rights Reserved | Register