Representation
More room for Filipino, Taglish, and the language we use in everyday life.
Why Makata exists
Language carries more than information. It carries how we think, the things we know, and the context that makes us understood. Philippine languages deserve a deeper place in AI research.
We want to help the Philippines build the capability to create, study, and operate AI—not only use it. That begins with patient, independent work, informed by the wider research community.
More room for Filipino, Taglish, and the language we use in everyday life.
Local knowledge and engineering skills to investigate the models behind the applications.
The long-term ability to make informed choices about our data, models, and infrastructure.
Research directions
Our initial focus is Filipino and Taglish. These are the areas we intend to investigate as our research capabilities develop.
Investigating how language models understand, generate, and reason in Filipino and Taglish, including the ways we move between languages.
Exploring the collection, curation, and documentation of Philippine-language data, with careful attention to provenance, privacy, and licensing.
Studying open-weight models and controlled adaptation experiments, while building the knowledge needed for original model research.
Planning reproducible evaluations that make capabilities, limitations, and performance in Philippine contexts visible.
Investigating smaller models and efficient approaches that could run on accessible hardware and locally controlled infrastructure.
Other Philippine languages are a longer-term direction, guided by suitable datasets and language expertise.
Our approach
Credibility comes from how the research is done—and what we are willing to make clear.
A name rooted in language, expression, and the knowledge we carry.
Document methods, configurations, and limitations so that others can understand and eventually repeat our experiments.
Treat linguistic nuance, code-switching, and cultural knowledge as central research questions.
Evaluate against clear baselines. Report regressions alongside improvements, and distinguish plans from demonstrated results.
Respect data rights and privacy. Be explicit about model origins and the difference between adaptation and training from scratch.
Where we are
Foundational stage
We are establishing Makata's research environment, documentation, and public identity.
Next comes defining manageable baseline experiments, evaluating candidate open-weight models, and investigating appropriately sourced data.
No Makata model has been trained or publicly released. Baseline evaluations and model adaptation are planned work.
An open invitation
Researchers, language specialists, engineers, universities, and institutions: if these questions matter to you, we'd like to connect.
Explore a collaboration jansencadorna5@gmail.com