Explainable header-centric framework maps 120,000 columns to data quality issues and semantic types
AI-summarised brief · reviewed before publication
Researchers unveiled a header‑centric framework that classifies spreadsheet column headers into 39 interpretable semantic types without inspecting cell values. The system records token‑level traceability, linking each classification to specific header words. Assigned types trigger validation rules that detect data quality issues such as missing values, duplicates, domain violations, and type mismatches. Results from about 120,000 headers across UCI, Kaggle, VizNet, and SemTab 2024 benchmarks show modest strict evaluation scores, but diagnostic audits reveal that many discrepancies arise from benchmark granularity and ontology choices rather than model failure. The framework also offers parallel knowledge‑graph mapping to DBpedia, enabling end‑to‑end data preparation for analytics pipelines.
💡 Why It Matters
- · By eliminating the need to access raw data, the approach speeds up data onboarding and preserves privacy, while its transparent decision trail aids auditability in regulated environments.