Public web sources turned into clean CSV/JSON datasets, every field and source documented. Strong on Chinese- and Southeast-Asian-language sources.