Open Source Datasets

Curated datasets sourced from government records, community discourse, and literary archives, vetted specifically for training.

Data Vetting & Linguistic Alignment

We don’t just stockpile data; we curate it. Our focus is on data vetting and linguistic alignment. By sourcing from government records, community discourse, and literary archives, and applying automated filters to strip away non-local phrasing, we ensure our models learn from authentic Taiwanese context rather than generic linguistic imports.
Illustration showing correct Traditional Chinese terminology

Formosa Vision

In collaboration with the Taiwan Cultural Memory Bank, we are building vision datasets that teach AI to decode local cultural symbols — from historic trails and military villages to the landscapes of Matsu. It’s about teaching the model to see the world through a local lens
Taiwanese traditional architecture illustration