Where Education Students Find Data: The 2026 Guide to US Education Datasets for a Dissertation, Study by Study

·

Where Education Students Find Data: The 2026 Guide to US Education Datasets for a Dissertation, Study by Study

Most education dissertations in the United States never collect a single survey. They analyze data that the federal government, a state agency or a university research center already collected from thousands of students, teachers and schools, and the quality of the dissertation depends on whether the student picked the right file. This guide profiles the datasets US education dissertations actually use, organized by the unit they measure, with the sample, the years, the variables, the access tier and the dissertation questions each one can carry. It complements our discipline-by-discipline directory of free datasets and open data repositories, which names the portals; this guide is about choosing a study.

Three questions before you open a data portal

What is your unit of analysis? Students, teachers, schools, districts or institutions. A question about student growth needs a student-level longitudinal file; a question about district spending needs the Common Core of Data, not a student survey. Does the question need change over time? Cross-sectional files such as NAEP answer how groups differ at one point; longitudinal studies such as ECLS-K:2011 and HSLS:09 follow the same students and answer how earlier conditions predict later outcomes. What access tier can you reach? Public-use files download in minutes; data-use agreements take days; a restricted-use license is held by your institution, not by you, and takes months. Choose the study with all three answers in hand, then write the research question in its final form.

Longitudinal student studies from NCES and BLS

  1. Early Childhood Longitudinal Study, Kindergarten Class of 2010-11 (ECLS-K:2011). A nationally representative sample of children who started kindergarten in 2010-11 in public and private schools, followed through fifth grade. The final public-use file holds nine rounds of data (fall 2010, spring 2011, fall 2011, spring 2012, fall 2012, spring 2013, spring 2014, spring 2015 and spring 2016), with 18,174 child records and roughly 26,100 variables covering direct child assessments in reading, mathematics and science, parent interviews, teacher and school administrator questionnaires, and measures of social, emotional and physical development. Dissertation fit: early achievement gaps, full-day versus part-day kindergarten, executive function and later reading, retention in grade, home and classroom predictors of growth. Access: public-use file with the Electronic Codebook for SAS, SPSS and Stata syntax; a restricted-use version adds finer geography and school identifiers.
  2. ECLS-K (1998-99) and ECLS-B. The original kindergarten cohort runs from kindergarten through eighth grade and remains the study for questions that need middle-school outcomes. The birth cohort, ECLS-B, followed children from nine months to kindergarten entry, and its case-level data are available only on restricted-use files because of the sensitivity of the data and the young age of the children.
  3. High School Longitudinal Study of 2009 (HSLS:09). More than 23,000 ninth graders in 944 schools, first surveyed in fall 2009, followed up in 2012, updated in 2013 with high school transcripts collected in 2013-14, followed again in 2016 and supplemented with postsecondary transcripts in 2017-18. Its distinctive content is the mathematics and science focus: algebra assessment, STEM course-taking, teacher and counselor questionnaires, and postsecondary choice. Dissertation fit: STEM pipeline questions, course-taking and college enrollment, counselor effects, first-generation college transitions. Access: public-use file for most variables; geocodes and some transcript detail require the restricted-use license.
  4. Education Longitudinal Study of 2002 (ELS:2002). A cohort of high school sophomores in 2002 followed through 2012, with transcripts, postsecondary enrollment and early labor market outcomes. Older than HSLS:09, but the only NCES high school cohort with a full decade of post-high-school outcomes, which makes it the file for questions about degree completion and early earnings.
  5. National Longitudinal Survey of Youth 1997 (NLSY97). Run by the Bureau of Labor Statistics, it follows a cohort born between 1980 and 1984 who were first interviewed in 1997 and are now in their forties, with schooling histories, high school transcripts, an aptitude test battery, employment, family formation and income. Education dissertations use it for the long-run returns to schooling, the timing of degree completion and the link between adolescent circumstances and adult work. Access: public data through the NLS Investigator after free registration; geocode files are restricted.
  6. Add Health. The National Longitudinal Study of Adolescent to Adult Health began with adolescents in grades 7 to 12 in 1994-95 and has followed them into their forties. Its school-level design, with in-school surveys and friendship networks, supports dissertations on school climate, peer effects and the health consequences of schooling. Public-use samples are available through ICPSR; the full sample and the network data require a contractual agreement.

Assessment data: NAEP, the international studies and SEDA

  1. National Assessment of Educational Progress (NAEP). The Nation’s Report Card assesses samples of students in grades 4, 8 and 12 in mathematics, reading, science, writing, U.S. history, civics, geography, economics, the arts and technology and engineering literacy, with results reported for the nation, states and participating urban districts. The NAEP Data Explorer produces weighted tables and significance tests online without a download. Student-level files are restricted-use, and scores are released as plausible values that must be analyzed as a set. Dissertation fit: state policy comparisons, achievement gaps over time, questionnaire correlates of performance.
  2. PISA, TIMSS and PIRLS. The international assessments include a US sample and publish microdata for free download. PISA tests 15-year-olds in reading, mathematics and science every three years under the OECD; TIMSS covers fourth and eighth grade mathematics and science; PIRLS covers fourth grade reading. Each carries student, teacher and school questionnaires and, like NAEP, plausible values and replicate weights. Dissertation fit: comparative questions, and single-country questions where the US NAEP student file is out of reach.
  3. Stanford Education Data Archive (SEDA). Built by the Educational Opportunity Project at Stanford, SEDA converts state test results into comparable achievement estimates for schools, districts, counties and states in grades 3 to 8 in English language arts and mathematics, with subgroup estimates by race, gender and economic disadvantage. Version 5.0 covers 2009 to 2019 (Reardon, Ho, Shear, Fahle, Kalogrides and Saliba, 2024), and the trends project extends the series through the pandemic years. Download is free after accepting a data-use agreement that bars commercial use and requires the archive to be cited as the source. Dissertation fit: district-level achievement gaps, segregation and opportunity, policy evaluation across states with a common metric.
Concentric layers showing student figures inside a classroom inside a school outline, with a clustered-dots chart beside them, illustrating nested education data
Student files, school files and district files answer different questions; the unit of analysis decides which dataset you need.

School, district and institution files

  1. Common Core of Data (CCD). The annual national census of all public elementary and secondary schools and school districts, with enrollment by grade and demographic group, free and reduced-price lunch eligibility, staffing counts, locale codes and, through the school district finance survey, revenues and expenditures. Access runs through the ElSi table generator, the school and district locators and the CCD data file tool for full downloads. Dissertation fit: any study that needs school or district characteristics as variables, and the merge key for almost every other dataset on this list.
  2. Civil Rights Data Collection (CRDC). Collected by the Department of Education’s Office for Civil Rights from public schools and districts nationwide on a two-year cycle, it reports suspensions and expulsions, referrals to law enforcement, enrollment in advanced coursework and gifted programs, restraint and seclusion, chronic absenteeism, teacher experience and school-level offerings, all disaggregated by race, sex, disability and English learner status. Dissertation fit: discipline disparities, access to advanced coursework, special education placement, the effects of school policing. Public files download from the CRDC site and merge to the CCD on the school identifier.
  3. Integrated Postsecondary Education Data System (IPEDS). The mandatory annual reporting system for more than 6,000 colleges, universities and technical and vocational institutions, with survey components on enrollment, completions, graduation rates, employees and staff, institutional revenues and financial aid. Complete data files are available from the 1980-81 collection year onward in CSV format, and the Data Explorer, custom data files and statistical tables handle smaller extracts. Dissertation fit: institutional comparisons, completion and retention rates, tuition and aid, faculty composition, the effect of state funding on outcomes.
  4. State longitudinal data systems and state report cards. Every state publishes school report card data, and most maintain a longitudinal data system that links students across years. Student-level state data are obtained through a data request to the state education agency or a research partnership, with a data-sharing agreement and often a fee; expect the process to take a semester. For a district-based EdD study this is frequently the only route to the outcomes that matter locally.

Postsecondary student studies

  1. NPSAS, BPS and B&B. The National Postsecondary Student Aid Study is a cross-sectional survey of how students pay for college; Beginning Postsecondary Students follows first-time students from NPSAS for six years; Baccalaureate and Beyond follows bachelor’s degree recipients into work and graduate school. Student-level files are restricted-use, but NCES DataLab (PowerStats) produces weighted tables and regressions from them online without a license.
  2. National Survey of Student Engagement (NSSE). Institutional data collected by Indiana University from participating colleges; individual researchers obtain it through their own institution’s assessment office rather than from a public file.

Which dataset fits which dissertation question

Matching common education dissertation questions to a dataset
Dissertation question Dataset Unit Access tier
Does full-day kindergarten predict reading growth through third grade? ECLS-K:2011 Student, longitudinal Public-use
Which ninth-grade factors predict a STEM major? HSLS:09 Student, longitudinal Public-use; geocodes restricted
Have racial gaps in grade 8 mathematics narrowed since 2003? NAEP Student samples, cross-sectional Data Explorer public; microdata restricted
Are district achievement gaps related to segregation? SEDA merged to CCD District Data-use agreement
Are suspension rates related to school resource officers? CRDC merged to CCD School Public-use
Does state appropriation per student predict graduation rates? IPEDS Institution, panel Public-use
How do first-generation students finance a degree? NPSAS or BPS Student DataLab public; microdata restricted
What are the long-run earnings returns to a two-year degree? NLSY97 or ELS:2002 Individual, longitudinal Public-use

Access tiers, and what a restricted-use license really involves

Three tiers cover almost everything above. Public-use files have been altered to prevent identification: small cells are suppressed, geography is coarsened, some variables are removed. They download after a click and are sufficient for most dissertations. Data-use agreements, as with SEDA or Add Health’s public sample, ask you to accept terms on use and citation and sometimes to register an email; they take minutes to days. Restricted-use licenses for NCES microdata are issued to institutions, not to individuals: the university signs, a principal project officer is named, every person who will touch the data signs an affidavit of nondisclosure, and a security plan describes the stand-alone computer the data will live on. A doctoral student is added to a faculty member’s license rather than holding one alone, and any analysis output leaving the licensed computer must meet disclosure rules. Budget three to six months and ask whether the public-use file or DataLab would answer the question first; very often it does. The same logic governs the safeguarded and controlled tiers described in our guide to UK longitudinal cohort studies, which is the closest parallel for students comparing the two systems.

Four obligations once you have the file

  1. Use the weights. Every NCES sample is stratified and clustered, and assessment scores come as plausible values. Our guide to which statistical test an education dissertation should use explains the replicate-weight and plausible-value procedures; a regression run without them will be sent back.
  2. Get the ethics determination in writing. Most public-use secondary analyses are exempt or not human subjects research, but the determination belongs to your IRB, not to you; our FAQ on whether your dissertation needs ethical approval walks through the categories.
  3. Document the variables you built. Recoded categories, merged files and derived scores need a data dictionary and a cleaning log, the practice set out in our guide to building a data dictionary for your thesis.
  4. Cite the file, not the agency. Reference the specific study, wave, file version and publication number, following the conventions in our guide to citing datasets, software and code in APA 7. Committees notice a reference list that cites NCES as if it were one document.

Turn the dataset into a methodology chapter

Tell Tesify which study you chose, the waves and variables you will use and the access tier you obtained, and it drafts the data-source section of your methodology: sampling design, weights, variable construction, the ethics determination and the limitations of secondary analysis, with every study documentation reference formatted in APA 7 by Auto Bibliography.

Start your education dissertation with Tesify, free to begin

Frequently asked questions

Can a doctoral student get an NCES restricted-use license?

Not alone. Licenses are issued to institutions, and a student is added as an authorized user under a faculty member’s license after signing an affidavit of nondisclosure. Ask whether the public-use file or DataLab answers the question before starting the application.

What is the difference between NAEP and SEDA?

NAEP is a sample-based federal assessment that reports for the nation, states and large districts, with student-level data restricted. SEDA is a Stanford project that places state test results on a common scale and reports estimates for every district, county and most schools, with download under a data-use agreement.

Which dataset should I use for a district-level EdD study?

Start with the Common Core of Data for school and district characteristics, add the Civil Rights Data Collection for discipline and course access, and SEDA for achievement. For student-level outcomes in your own district, request the data from the district or the state longitudinal data system under a data-sharing agreement.

Do I need IRB approval to analyze public-use NCES data?

Usually the IRB will determine that the study is not human subjects research or is exempt because the data are de-identified, but you must submit the determination request and keep the letter. Restricted-use data typically require a full review because the files contain finer identifiers.

Which study has the most recent data on kindergarten children?

ECLS-K:2011 is the most recent kindergarten cohort with released public-use data, followed through fifth grade in 2016. A new kindergarten cohort, ECLS-K:2024, is in the field, and its files will appear over the coming years.

Can I merge these datasets?

Yes, at the school or district level through the NCES school and agency identifiers used by the CCD, the CRDC and SEDA, and at the institution level through the IPEDS unit identifier. Student-level merges across studies are not possible because the samples differ.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *