TruongThuHa GitHub avatar

diem_thi_thptqg-2026-

TruongThuHa

A dataset of 1,200,202 Vietnamese high school exam scores from 34 provinces, scraped from a public source, with per-province CSV files and statistical anomaly analyses.

Stars

8

7-day growth

No data

Forks

0

Open issues

0

License

No data

Last updated

2026-07-04

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides a large-scale, well-structured public dataset of individual exam scores, includes ready-to-use scraping scripts, and contains detailed anomaly detection (e.g., unusual clustering of scores at 5.0 in literature, high math score clusters).

Who it is for

  • Education researchers and policy analysts
  • Data scientists and statisticians
  • Vietnamese developers building education tools
  • Journalists covering education performance

Use cases

  • Analyzing provincial exam performance and trends
  • Detecting and visualizing score distribution anomalies
  • Benchmarking school or district performance
  • Building dashboards for educational monitoring

Strengths

  • Covers 1.2 million exam takers across 34 reorganized provinces
  • Clean CSV structure with per-province files and consistent column format
  • Includes async scraping scripts (Python + aiohttp) for reproducibility
  • Provides data-driven anomaly analysis with concrete numbers and tables

Considerations

  • Contains personally identifiable information (SBD numbers) – only for personal research, not for redistribution or commercial use
  • May have missing records due to discontinuous SBD numbers or post-collection edits
  • Anomaly analyses are statistical observations, not official conclusions

README quick start

Điểm thi THPT Quốc gia 2026

Bộ dữ liệu điểm thi tốt nghiệp THPT năm 2026 của 1.200.202 thí sinh toàn quốc (34 tỉnh/thành phố sau khi sắp xếp lại đơn vị hành chính).

Dữ liệu thu thập từ trang tra cứu công khai của VietnamNet ngày 01/07/2026.


Cấu trúc dữ liệu

Thư mục data/

FileMô tả
diem_thi_thptqg_2026_all.csvToàn quốc, đã loại trùng, sắp xếp theo SBD
01-ha-noi.csvHà Nội
04-cao-bang.csvCao Bằng
08-tuyen-quang.csvTuyên Quang
11-dien-bien.csvĐiện Biên
12-lai-chau.csvLai Châu
14-son-la.csvSơn La
15-lao-cai.csvLào Cai
19-thai-nguyen.csvThái Nguyên
20-lang-son.csvLạng Sơn
22-quang-ninh.csvQuảng Ninh
24-bac-ninh.csvBắc Ninh
25-phu-tho.csvPhú Thọ
31-hai-phong.csvHải Phòng
33-hung-yen.csvHưng Yên
37-ninh-binh.csvNinh Bình
38-thanh-hoa.csvThanh Hóa
40-nghe-an.csvNghệ An
42-ha-tinh.csvHà Tĩnh
44-quang-tri.csvQuảng Trị
46-hue.csvHuế
48-da-nang.csvĐà Nẵng
51-quang-ngai.csvQuảng Ngãi
52-gia-lai.csvGia Lai
56-khanh-hoa.csvKhánh Hòa
66-dak-lak.csvĐắk Lắk
68-lam-dong.csvLâm Đồng
75-dong-nai.csvĐồng Nai
79-ho-chi-minh.csvHồ Chí Minh
80-tay-ninh.csvTây Ninh
82-dong-thap.csvĐồng Tháp
86-vinh-long.csvVĩnh Long
91-an-giang.csvAn Giang
92-can-tho.csvCần Thơ
96-ca-mau.csvCà Mau

Cột CSV

CộtÝ nghĩa
sbdSố báo danh (8 chữ số: 2 mã tỉnh + 6 số thứ tự)
province_codeMã tỉnh (2 chữ số đầu của SBD)
province_nameTên tỉnh/thành phố
toanToán
ngu_vanNgữ văn
ngoai_nguNgoại ngữ
vat_liVật lý
hoa_hocHóa học
sinh_hocSinh học
lich_suLịch sử
dia_liĐịa lý
gdcdGiáo dục công dân
tin_hocTin học
cong_ngheCông nghệ
extraMôn khác (nếu có)

Điểm còn trống ("") = thí sinh không thi môn đó.


Ví dụ sử dụng

Python / pandas

import pandas as pd

# Toàn quốc
df = pd.read_csv("data/diem_thi_thptqg_2026_all.csv", dtype={"sbd": str, "province_code": str})

# Một tỉnh
hn = pd.read_csv("data/01-ha-noi.csv", dtype={"sbd": str})

# Điểm trung bình Toán theo tỉnh
print(df.groupby

Description

Bộ dữ liệu điểm thi tốt nghiệp THPT năm 2026 của 1.200.202 thí sinh toàn quốc (34 tỉnh/thành phố sau khi sắp xếp lại đơn vị hành chính). Dữ liệu thu thập từ trang tra cứu công khai của VietnamNet ngày 01/07/2026.

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482