Back to events
Recording · Field notes
Multilingual retrieval for Southeast Asian products
What breaks when your corpus is half Bahasa Malaysia, a third English, and the rest code-switched, and how to measure it.
Benchmarks are English. Your customers are not. We walk through the specific failure modes of retrieval on code-switched Malaysian text, the embedding choices that matter, and how to build an eval set that reflects the language your users actually type.
Speakers
ML Engineer, Enterprise
Retrieval and evaluation
Start before the gap gets expensive.
Tell us where the work is stuck. We will tell you, in plain language, what AI can and cannot fix, and what it would take to do it properly.