Continuous localization for Vietnamese: a Crowdin setup guide
💡 Continuous localization keeps Vietnamese strings in sync with every release by automating the pull from your source repo into Crowdin and pushing approved translations back as a pull request. The core workflow: connect GitHub, add a
crowdin.ymlin your repo root, and set an approval threshold for Vietnamese so only human-reviewed strings merge. The main Vietnamese trap is Unicode normalization drift, where unchanged strings appear new and waste translation budget.
Key takeaways
- Crowdin's GitHub integration syncs hourly by default: new source strings upload automatically, approved Vietnamese translations come back as an
l10n_branch pull request. - A four-field
crowdin.yml(project_id, api_token, source path, translation path with%locale%) is all you need to start; the CLI skips unchanged files to keep CI runs fast. - Normalize source JSON to NFC Unicode before each push to avoid false "new string" detection caused by macOS NFD encoding.
- Vietnamese runs 10-25% longer than English; run automated layout tests with the
vilocale after every significant release. - Require human approval for Vietnamese strings before they merge - MT handles general UI copy well, but tone-mark errors in customer-facing text go undetected by automated quality estimation.
What is continuous localization, and why does Vietnamese lag behind?
Traditional localization follows a waterfall: developers freeze strings, hand them to translators, wait for the batch, then ship. In a modern agile product, this means Vietnamese ships weeks behind English - a visible trust gap for Vietnam's 78 million internet users (DataReportal, 2024).
Continuous localization replaces the string freeze with an automated trigger. The moment a developer pushes new strings to the source repository, the TMS detects the diff and routes it to translators. When they approve, the platform opens a pull request and Vietnamese ships with the same release. No batch, no lag.
The result: Vietnamese and English launch together, localization costs stay predictable because translators handle small batches continuously rather than large dumps quarterly, and the translation memory grows sprint by sprint.
How does the Crowdin and GitHub sync work?
Crowdin's GitHub integration runs bidirectional. When you push your source file (for example src/locales/en.json) to GitHub, Crowdin pulls the new or changed strings into the project. When translators approve the Vietnamese strings, the platform pushes the translated file to an l10n_ prefixed branch and opens a pull request automatically.
The sync interval is hourly by default. You can trigger an immediate sync with Sync Now in the Crowdin dashboard, or use the CLI in a CI/CD pipeline for finer control. The integration also supports branch management: separate Crowdin branches for main, staging, and feature branches keep translation state isolated per release track.
Alternative TMS platforms - Phrase TMS, Lokalise, and Smartcat - offer the same webhook-triggered pattern. The core concepts (source file upload on push, translation PR on approval) are consistent; they differ mainly in pricing, workflow automation depth, and MT integration options.
What goes in a minimal crowdin.yml for Vietnamese?
Place a crowdin.yml file in your repository root. The minimum viable configuration for a JSON-based Vietnamese project:
project_id: "YOUR_PROJECT_ID"
api_token: "%CROWDIN_API_TOKEN%"
base_path: "."
files:
- source: /src/locales/en.json
translation: /src/locales/%locale%.json
ignore:
- /src/locales/en.json
The %locale% placeholder maps to vi for Vietnamese, producing src/locales/vi.json. The ignore rule prevents the source file from being overwritten as a translation target.
Add two steps to your CI pipeline (GitHub Actions example):
- name: Upload sources to Crowdin
run: crowdin upload sources --config crowdin.yml
- name: Download approved Vietnamese
run: crowdin download --language vi --config crowdin.yml
Store CROWDIN_API_TOKEN as a repository secret. The Crowdin CLI uses cached source uploads and skips files unchanged since the last run, so CI stays fast even with hundreds of strings.
What Vietnamese-specific pitfalls break automated workflows?
Three failure modes are unique to Vietnamese in continuous localization pipelines:
Unicode normalization drift. Vietnamese diacritics have two valid Unicode forms: precomposed NFC (one code point per accented character) and decomposed NFD (base letter plus combining mark). macOS editors often write NFD; most TMS platforms store NFC. When your source file arrives in NFD, the translation memory sees every diacritic-bearing string as new, match scores drop to zero, and you pay for retranslation. Add .normalize("NFC") to your source files in a pre-commit hook.
Tone-mark errors. Vietnamese has six tones written as diacritics. A wrong tone changes the word entirely: "ma" (ghost), "mà" (but), "má" (cheek), "mả" (tomb), "mã" (code), "mạ" (rice seedling). Automated quality estimation tools do not flag tone errors because the output is still valid Vietnamese text. Require human review for any customer-facing strings.
String length growth. Vietnamese strings commonly run 10-25% longer than English. Navigation labels, button text, and mobile tab titles overflow their containers silently in automated builds. Run layout regression tests with the vi locale to catch overflow before users do.
What your team can handle in-house vs when a native Vietnamese specialist saves money
The automation layer - Crowdin setup, crowdin.yml, GitHub integration, CI/CD pipeline - is a developer task. No specialist needed.
For the translation layer, the practical dividing line is:
- General UI copy and informational content: MT with a native Vietnamese review pass is cost-effective. Modern neural MT handles common product vocabulary well; the review pass catches tone errors and unnatural phrasing before users see them.
- Marketing and brand-sensitive copy: A native translator who knows Northern Vietnamese register (the standard for formal business) and your product voice produces copy that converts. Generic MT output reads as literally translated to Vietnamese users.
- Technical, legal, or financial strings: Domain-specific Sino-Vietnamese terminology matters. A fintech app that uses "tài khoản giao dịch" where "tài khoản thanh toán" is expected signals an outsider to Vietnamese users. A native specialist chooses the term users actually use.
The Vietnamese website localization service covers both the pipeline review and the human translation layer for web products entering Vietnam.
FAQ
Can I use Crowdin's free plan for continuous Vietnamese localization?
Crowdin's free tier supports one active project with limited features. The GitHub integration and branch management that make continuous localization practical require a paid Crowdin plan. Phrase TMS and Lokalise offer comparable paid plans; compare current rates on their respective sites.
What locale code should I use for Vietnamese in my locale files?
Use vi (ISO 639-1) for Vietnamese without regional variant, or vi-VN (IETF BCP 47) when you need to specify Vietnam as the region. In Crowdin's crowdin.yml, the %locale% placeholder outputs the locale code as configured in your project settings.
Does Crowdin support Vietnamese translation memory and glossary?
Yes. Crowdin maintains a project-level translation memory for Vietnamese. When a source string changes, the TM surfaces the previous Vietnamese translation with a match score, reducing retranslation cost. Glossaries let you lock Sino-Vietnamese technical terms so multiple translators use consistent vocabulary.
How do I stop machine-translated Vietnamese from shipping to production?
Set a minimum approval threshold in Crowdin's project settings and configure the GitHub integration to create a translation PR only after all strings reach that threshold. You can also add the --export-approved-only flag to crowdin download in your CI pipeline to download only human-approved strings.
Does continuous localization work for mobile apps as well as websites?
Yes. Crowdin supports Android XML, iOS .strings, and Flutter ARB file formats alongside JSON. The GitHub integration works identically: source strings push to Crowdin on every commit, approved Vietnamese translations come back as a PR. The same Unicode normalization and layout testing advice applies to mobile.
Official Sources
- Crowdin GitHub Integration Documentation - branch naming convention, sync interval, crowdin.yml config, PR workflow. Verified Oct 2026.
- Crowdin CLI Documentation - upload/download commands, CI/CD integration, configuration keys. Verified Oct 2026.
- Phrase: Continuous Localization Guide - workflow principles and TMS/VCS integration patterns. Verified Oct 2026.
- W3C: Localization vs. Internationalization - authoritative definitions of i18n and l10n. Verified Oct 2026.
- DataReportal: Digital 2024 Vietnam - Vietnam internet user statistics used in this article. Verified Oct 2026.
Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →
