Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
38 commits
Select commit Hold shift + click to select a range
0afe7b5
Updated code
smarthg-gi Aug 23, 2026
a446adb
Replaced lambda with def _PV_FORMAT(pv)
smarthg-gi Aug 24, 2026
23cefa0
updated manifest with nodes.mcf
smarthg-gi Aug 24, 2026
cf7777b
Adressed CRA comments
smarthg-gi Aug 24, 2026
c56f916
Made changes as per CRA review
smarthg-gi Aug 24, 2026
bcdd014
Updated test_data
smarthg-gi Aug 24, 2026
344e705
Adding files for PublicSchool
smarthg-gi Aug 30, 2026
e3139ba
Merge branch 'master' into NCES_SchoolDistrict_fix
smarthg-gi Aug 31, 2026
6d4840e
Fix Python formatting as per yapf style
smarthg-gi Aug 31, 2026
a2c33da
Updated place merge logic
smarthg-gi Sep 2, 2026
132eccf
Resolve merge conflict: keep simplified multi-file place ingestion
smarthg-gi Sep 2, 2026
a815f56
Fix Python formatting as per yapf style
smarthg-gi Sep 2, 2026
d2eae4a
updated test files
smarthg-gi Sep 2, 2026
eed877d
Update provenance_url to https://nces.ed.gov/ in public_school and sc…
smarthg-gi Sep 2, 2026
e40737b
Merge branch 'datacommonsorg:master' into NCES_SchoolDistrict_fix
smarthg-gi Sep 6, 2026
670cce9
updated files as per review comments
smarthg-gi Sep 7, 2026
9111c51
Fix district ID padding and address review comments
smarthg-gi Sep 7, 2026
31bc976
Fix review findings for error handling, test hermeticity, and manifes…
smarthg-gi Sep 9, 2026
de3ecf2
removing redundant files
smarthg-gi Sep 9, 2026
4f704c7
Regenerate test output files using generator scripts and align test m…
smarthg-gi Sep 9, 2026
b50e85e
adding data freshness and max date consistency rules
smarthg-gi Sep 9, 2026
d968067
use absl.app.run(main) in process.py
smarthg-gi Sep 9, 2026
46e77e2
added validation rules for public school
smarthg-gi Sep 10, 2026
6604e27
Fix python formatting as per yapf style and trim extra newlines in va…
smarthg-gi Sep 10, 2026
b48e4d4
Used raw string literals for regex in us_education.py to resolve Pyth…
smarthg-gi Sep 10, 2026
7cd4223
Fix duplicate function call in lambda and clean up unused imports
smarthg-gi Sep 11, 2026
fd04f16
updated files with the split of imports
smarthg-gi Oct 4, 2026
a7d9844
remove validation_config from place imports and fail fast on missing …
smarthg-gi Oct 4, 2026
c17f846
Resolve PR review comments: dynamic footer parsing, address cleanup, …
smarthg-gi Oct 5, 2026
c9b2301
Fix School_Type_Public educationalMethod mapping, dcs prefix, and REA…
smarthg-gi Oct 5, 2026
ad9758f
fix cross-year attribute bleeding and preserve unreadable sentinels i…
smarthg-gi Oct 5, 2026
1c5b40c
Address review feedback: place coalescing, sentinel mappings, config …
smarthg-gi Oct 6, 2026
58a5d76
fix yapf formatting
smarthg-gi Oct 6, 2026
38ba5da
fix yapf formatting
smarthg-gi Oct 6, 2026
3477eae
Resolve PR comments: fix place enum prefixes, sentinel filtering, and…
smarthg-gi Oct 6, 2026
776fffd
Remove dcs prefix from School_Management to resolve missing reference…
smarthg-gi Oct 6, 2026
354f4aa
updated correct config
smarthg-gi Oct 7, 2026
df7226f
updated cron scheduls for timely execution
smarthg-gi Oct 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
60 changes: 54 additions & 6 deletions scripts/us_nces/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@
## Import Overview:
This dataset has Population Estimates for the National Center for Education Statistics in US for
- Private School - 1997-98 to 2019-20
- School District -2010-11 to 2023-24
- Public Schools - 2010-11 to 2023-24
- School District - 2010-11 to 2024-25
- Public Schools - 2010-11 to 2024-25

## Source URL:
https://nces.ed.gov/ccd/elsi/tableGenerator.aspx
Expand Down Expand Up @@ -34,9 +34,57 @@ This dataset has Population Estimates for the National Center for Education Stat
The only manual part here is after downloading the input files and then uploading them to gcp bucket. Once they're uploaded, Each import requires its own sh command to copy the files from Google Cloud to a local folder called gcs_folder/input_files. From there, a script automatically picks up these files to process them. Finally, it generates the output and saves it in gcs_folder/output_files

### Script Execution Details
private : python3 private_school/process.py
public : python3 public_school/process.py
district : python3 school_district/process.py
Each domain's `process.py` supports the `--mode` flag (`place`, `stats`, or `all`, default: `all`):

```bash
# Public School
python3 public_school/process.py --mode=place # Place data only
python3 public_school/process.py --mode=stats # Statistical observations only
python3 public_school/process.py --mode=all # Both place and stats (default)

# School District
python3 school_district/process.py --mode=place
python3 school_district/process.py --mode=stats
python3 school_district/process.py --mode=all

# Private School
python3 private_school/process.py --mode=place
Comment thread
smarthg-gi marked this conversation as resolved.
python3 private_school/process.py --mode=stats
python3 private_school/process.py --mode=all
```

### Import Architecture & Staggered Cron Scheduling
Each domain is split into independent Cloud Batch imports in `manifest.json`. Because newer NCES data introduces newly defined places (such as brand-new school districts and public schools) that are referenced across imports (e.g., `NCES_PublicSchool` places reference parent school districts via `schoolDistrict: dcs:geoId/sch...`, and Stats imports reference places via `observationAbout`), jobs are organized into **three staggered tiers spaced 1 week (7 days) apart**.

This 1-week buffer ensures that Data Commons Production ingests and indexes the newly defined Place nodes from Tier $N$ before Tier $N+1$ jobs execute, guaranteeing zero missing reference warnings (`check_missing_refs_count: PASSED`).

#### Execution Tiers & Dependencies

```text
[TIER 1: Week 1 (Day 3)] [TIER 2: Week 2 (Day 10, +7d)] [TIER 3: Week 3 (Day 17, +14d)]
======================== ============================== ===============================

NCES_SchoolDistrict (Place) ─────────► NCES_SchoolDistrictStats (Stats)
(Defines: dcid:geoId/sch...) │ (Refs: observationAbout: dcs:geoId/sch...)
│
└───► NCES_PublicSchool (Place) ───────────────► NCES_PublicSchoolStats (Stats)
(Defines: dcid:nces/... [Public]) (Refs: observationAbout: dcs:nces/...)
(Refs: schoolDistrict: dcs:geoId/sch...)

NCES_PrivateSchool (Place) ──────────► NCES_PrivateSchoolStats (Stats)
(Defines: dcid:nces/... [Private]) (Refs: observationAbout: dcs:nces/...)
```

#### Staggered Quarterly Cron Schedule (`manifest.json`)

| Execution Tier | Import Name | Domain | Manifest Path | Staggered `cron_schedule` |
| :--- | :--- | :--- | :--- | :--- |
| **Tier 1 (Week 1, Day 3)** | `NCES_SchoolDistrict` | School District (Place) | `school_district/manifest.json` | `"30 3 3 3,6,9,12 *"` |
| **Tier 1 (Week 1, Day 3)** | `NCES_PrivateSchool` | Private School (Place) | `private_school/manifest.json` | `"30 4 3 3,6,9,12 *"` |
| **Tier 2 (Week 2, Day 10)** | `NCES_SchoolDistrictStats` | School District (Stats) | `school_district/manifest.json` | `"30 3 10 3,6,9,12 *"` |
| **Tier 2 (Week 2, Day 10)** | `NCES_PublicSchool` | Public School (Place) | `public_school/manifest.json` | `"30 5 10 3,6,9,12 *"` |
| **Tier 2 (Week 2, Day 10)** | `NCES_PrivateSchoolStats` | Private School (Stats)* | `private_school/manifest.json` | `"30 7 10 3,6,9,12 *"` |
| **Tier 3 (Week 3, Day 17)** | `NCES_PublicSchoolStats` | Public School (Stats) | `public_school/manifest.json` | `"30 3 17 3,6,9,12 *"` |


#### Cleaned Data
Expand Down Expand Up @@ -183,7 +231,7 @@ step 1 :

step 2 : Use the command-line tool to do genmcf using the CSV and TMCF files.

`java -jar '/usr/local/google/home/spateriya/Downloads/datacommons-import-tool-0.1-alpha.1-jar-with-dependencies.jar' genmcf -r FULL <place csv path> <place tmcf path>`
`java -jar <path_to_datacommons_import_tool>/datacommons-import-tool-jar-with-dependencies.jar genmcf -r FULL <place csv path> <place tmcf path>`

step 3 : Update the file path in the textproto files for NCES_PrivateSchool, NCES_PublicSchool, and NCES_SchoolDistrict.

Expand Down
28 changes: 20 additions & 8 deletions scripts/us_nces/common/prop_conf.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@
While preprocessing files column names are changed to SV names as used in
DC import
"""
import pandas as pd

# TMCF template for Demographics data. It changes based on import name.
TMCF_TEMPLATE = (
"Node: E:us_nces_demographics_{import_name}->E0\n"
Expand Down Expand Up @@ -121,19 +123,29 @@
"LEA Administrative Support Staff": "Faculty",
"School Administrators": "Faculty",
"School Administrative Support Staff": "Faculty",
"Student Support Services Staff": "Faculty",
# Note: "Student Support Services Staff" is intentionally matched by
# r"Student" to preserve backward compatibility with the existing StatVar
# Count_Student_StudentSupportServicesStaff in us_nces_stat_vars.mcf:573.
"School Psychologist": "Faculty",
"Other Support Services Staff": "Faculty"
"Other Support Services Staff": "Faculty",
"Total Staff": "Faculty"
Comment thread
smarthg-gi marked this conversation as resolved.
}
# One specific column comes under school grade property.
_SCHOOL_GRADE_PROP = {"Ungraded Students": "NCESUngradedClasses"}
# melting the columns based on sv_name column.
MELT_VAR_COL = "sv_name"


def _PV_FORMAT(pv):
"""Formats property-value pairs for MCF nodes; modified based on column."""
t = tuple(pv)
val = str(t[1]).strip() if not pd.isna(t[1]) else ""
if not val or val in ('None', 'nan', '<NA>'):
return ""
return f'"{t[0]}": "dcs:{val}"'


# pylint:disable=unnecessary-lambda-assignment
# Creating property pattern and the pattern is modified if required based on column.
_PV_FORMAT = lambda prop_val: f'"{prop_val[0]}": "dcs:{prop_val[1]}"' \
if 'None' not in prop_val[1] else ""
_UPDATE_MEASUREMENT_DENO = lambda prop: _DENOMINATOR_PROP.get(prop, prop)
_UPDATE_POPULATION_TYPE = lambda prop: _POPULATION_PROP.get(prop, "Student")
_UPDATE_GRADE_LEVEL = lambda prop: _SCHOOL_GRADE_PROP.get(prop, prop)
Expand Down Expand Up @@ -167,7 +179,7 @@
r")")

_SCHOOL_GRADE_PATTERN = (r"("
r"Grade \d{,2}"
r"Grade \d{1,2}"
r"|"
r"Prekindergarten and Kindergarten"
r"|"
Expand Down Expand Up @@ -227,15 +239,15 @@
r"|"
r"LEA Administrative Support Staff"
r"|"
r"Student Support Services Staff"
r"|"
r"School Administrative Support Staff"
r"|"
r"School Administrators"
r"|"
r"School Psychologist"
r"|"
r"Other Support Services Staff"
r"|"
r"Total Staff"
r")")

_GENDER_PATTERN = (r"("
Expand Down
47 changes: 34 additions & 13 deletions scripts/us_nces/common/replacement_functions.py
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,8 @@
"NCES_DataMissing",
"–":
"NCES_DataMissing",
"†":
"NCES_DataNotApplicable",
}

_SCHOOL_GRADE = {
Expand Down Expand Up @@ -173,6 +175,7 @@
"Grade 13": "SchoolGrade13",
np.nan: "NCES_GradeDataMissing",
"–": "NCES_GradeDataMissing",
"†": "NCES_GradeDataNotApplicable",
"All Ungraded": "NCESUngradedClasses",
"Adult Education": "AdultEducation",
"Transitional 1st grade": "TransitionalGrade1",
Expand Down Expand Up @@ -231,11 +234,12 @@
"Grade 13": "SchoolGrade13",
np.nan: "NCES_GradeDataMissing",
"–": "NCES_GradeDataMissing",
"†": "NCES_GradeDataNotApplicable",
"Adult Education": "AdultEducation",
"Transitional 1st grade": "TransitionalGrade1"
}

_PHYSICAL_ADD = {np.nan: "", "–": "", "Po Box": "PO BOX"}
_PHYSICAL_ADD = {np.nan: "", "–": "", "†": "", "Po Box": "PO BOX"}

# pylint:disable=line-too-long
_SCHOOL_LEVEL = {
Expand Down Expand Up @@ -276,9 +280,11 @@
"Secondary":
"SecondarySchool",
np.nan:
"",
"NCES_SchoolLevelDataMissing",
"–":
""
"NCES_SchoolLevelDataMissing",
"†":
"NCES_SchoolLevelDataNotApplicable"
}
# pylint:enable=line-too-long

Expand All @@ -297,6 +303,7 @@
_LUNCH = {
"Reduced-price Lunch":
"ReducedLunch",
# Legacy mapping preserved for historical time-series continuity with NCES_PublicSchoolStats.
"Free and Reduced Lunch":
"DirectCertificationLunch",
"Free Lunch":
Expand All @@ -314,9 +321,11 @@
"Yes under Provision 3":
"NCES_NationalSchoolLunchProgramYesUnderProvision3",
np.nan:
"NCES_MagnetDataMissing",
"NCES_NationalSchoolLunchProgramDataMissing",
"–":
"NCES_MagnetDataMissing"
"NCES_NationalSchoolLunchProgramDataMissing",
"†":
"NCES_NationalSchoolLunchProgramDataNotApplicable"
}

_SCHOOL_STAFF = {
Expand Down Expand Up @@ -372,7 +381,7 @@
"Phone Number": "PhoneNumber"
}

_GENDER = {"female": "Female", "male": "Male"}
_GENDER = {r"\bfemale\b": "Female", r"\bmale\b": "Male"}

_LOCALE = {
'13-City: Small': "NCES_CitySmall",
Expand All @@ -388,34 +397,39 @@
'31-Town: Fringe': "NCES_TownFringe",
'43-Rural: Remote': "NCES_RuralRemote",
np.nan: "NCES_LocaleDataMissing",
"–": "NCES_LocaleDataMissing"
"–": "NCES_LocaleDataMissing",
"†": "NCES_LocaleDataNotApplicable"
}

_UNREADABLE_TEXT = {"–": np.nan, "†": np.nan}

_NAN = {np.nan: "", "nan": ""}
_NAN = {np.nan: "", "nan": "", "–": "", "†": ""}

_CITY = {np.nan: "", "–": "", " ": ""}
_CITY = {np.nan: "", "–": "", "†": "", " ": ""}

_MAGNET = {
"1-Yes": "NCES_MagnetYes",
"2-No": "NCES_MagnetNo",
np.nan: "NCES_MagnetDataMissing",
"–": "NCES_MagnetDataMissing"
"–": "NCES_MagnetDataMissing",
"†": "NCES_MagnetDataNotApplicable"
}

_CHARTER = {
np.nan: "NCES_CharterDataMissing",
"1-Yes": "NCES_CharterYes",
"2-No": "NCES_CharterNo",
"–": "NCES_CharterDataMissing"
"–": "NCES_CharterDataMissing",
"†": "NCES_CharterDataNotApplicable"
}

_TITLE = {
np.nan:
"NCES_TitleISchoolStatusDataMissing",
"–":
"NCES_TitleISchoolStatusDataMissing",
"†":
"NCES_TitleISchoolStatusDataNotApplicable",
"6-Not a Title I school":
"NCES_TitleISchoolStatusNotEligible",
"5-Title I schoolwide school":
Expand All @@ -432,11 +446,17 @@

_SCHOOL_PUBLIC_TYPE = {
"1-Regular school": "NCES_PublicSchoolTypeRegular",
"4-Alternative/other school": "NCES_PublicSchoolTypeOther",
"1-Regular School": "NCES_PublicSchoolTypeRegular",
"2-Special education school": "NCES_PublicSchoolTypeSpecialEducation",
"2-Special Education School": "NCES_PublicSchoolTypeSpecialEducation",
"3-Vocational school": "NCES_PublicSchoolTypeVocational",
"3-Career and Technical school": "NCES_PublicSchoolTypeVocational",
"3-Career and Technical School": "NCES_PublicSchoolTypeVocational",
"4-Alternative/other school": "NCES_PublicSchoolTypeOther",
"4-Alternative Education School": "NCES_PublicSchoolTypeOther",
np.nan: "NCES_PublicSchoolTypeDataMissing",
"–": "NCES_PublicSchoolTypeDataMissing"
"–": "NCES_PublicSchoolTypeDataMissing",
"†": "NCES_PublicSchoolTypeDataNotApplicable",
}

_STATE_NAME = {
Expand Down Expand Up @@ -501,6 +521,7 @@ def replace_values(data_df: pd.DataFrame,
"Agency_Name": _NAN,
"School_Level_17": _SCHOOL_LEVEL,
"School_Level_16": _SCHOOL_LEVEL,
"School_Level": _SCHOOL_LEVEL,
"State_Agency_ID": _NAN,
"State_School_ID": _NAN,
"State_Name": _STATE_NAME
Expand Down
Loading
Loading