Data for individual socio-economic status come from:
cen_individ : The original baseline survey (this table is combined with crs_memball to add crshse & crsno) [2001-2004]
crs_cenm : in-migration form [gathered SES data 2001-2005]
css_sei_all : Annual individual survey [since 2007]
NCD_combined_data_all_final [derived table] : NCD baseline survey [2016-2018]
combined_tb_cases_controls_final : Combined TB data (derived dataset) [gathered SES data on TB cases and controls 1998-2018]
Each individual dataset is processed to ensure harmonisation of variable names and coding. They are then appended together and some summary variables created. The resulting dataset has one line per person per date (i.e. if a person has taken part in several surveys/studies they will have multiple records). There has been no attempt to 'clean' the data, i.e. to check whether subsequent schooling reports makes sense etc: cleaned schooling episodes are available on the derived dataset 'cleaned_school_episodes'.
In the variable labels (S) indicates that the variable is unchanged from the original source and (D) indicates that it is a 'derived' variable, often using more than one source variable. For both types information on where the data come from and how they are processed is found in the 'Recoding and Derivation' section
| Cases: | 499495 |
| Variables: | 36 |