Hey everyone! 👋
I've been deep in the weeds lately trying to standardize how my team handles multi-language content for our documentation and marketing sites, and I wanted to share what I've learned—and hopefully get your insights too. We use a mix of static site generators and headless CMS setups, and the translation layer was always a bit of an afterthought until it became a scaling nightmare.
From my experience, the "best" way really depends on your stack and workflow, but I'm a firm believer in treating translations as a first-class citizen in your CI/CD pipeline. You can't just bolt it on at the end.
Here’s the approach that’s been working for us, broken down:
**1. Separate Content from Structure**
Keep your translatable strings in a format that's agnostic to your presentation layer. We use `JSON` or `YAML` files structured by locale.
```json
// locales/en.json
{
"welcome": {
"heading": "Welcome to StackInsight",
"subtitle": "Join the community"
}
}
// locales/es.json
{
"welcome": {
"heading": "Bienvenido a StackInsight",
"subtitle": "Únete a la comunidad"
}
}
```
**2. Automate the Translation Sync**
Manual copy-pasting is a trap. We use a CLI tool (`i18next-parser` for our JS projects, `ttxt` for others) to extract new keys from source code into our base language file. Then, we push *only* the new or modified keys to a translation service (like Crowdin or Lokalise) via their API. The updated translations are pulled back as part of our nightly build.
**3. Integrate into the Build Process**
This is the GitOps/devops part I love. Our build script checks for missing translations. If a required locale is incomplete, the build *fails*, not just falls back. This forces us to keep everything in sync.
```bash
#!/bin/bash
# Example build check snippet
for locale in en es fr de;
do
if ! jq empty "./locales/${locale}.json"; then
echo "❌ Invalid JSON for ${locale}"
exit 1
fi
done
echo "✅ All locale files are valid."
```
**4. Handle Dynamic Content Differently**
For user-generated content or content from a CMS, we use a database with a `translations` table linked by a common key. The API then selects the correct language based on the `Accept-Language` header or a user preference.
The biggest pitfalls we've hit:
* **Context for translators:** A string like "Run" could be a verb or a noun. We now always add comments/context in our translation files.
* **Layout breaks:** Translated text can be much longer or shorter. Design with flexibility in mind—use CSS that accommodates text expansion.
* **Not translating URLs:** Remember to localize slugs if needed (`/en/blog/post` vs `/es/blog/articulo`), and consider hreflang tags for SEO.
I'm curious—how are you all managing this? Are you using any specific services, or have you built an in-house solution? Any horror stories or brilliant automations to share?
— francesc
— francesc
Real translation engineer at a mid-market SaaS with 100 devs, running a Next.js i18n stack in production that we built in-house after a costly experiment with an enterprise platform.
1. **Platform Fit & Target:** The "full-service" platforms (Crowdin, Phrase) are built for non-technical localization managers in a 500+ person enterprise. For a dev team of under 50, they're organizational overkill. The cheaper ones (Weblate, Tolgee) assume you're an open-source project or a 5-person startup.
2. **Real Total Cost:** The advertised "starting at $XX/user/month" is a lie. For a team our size, the true cost of a platform like Crowdin was $45k/year after the mandatory "Collaboration" add-on and the per-word MT overages. Our in-house system (using Postgres, i18next, and a custom sync script) runs at under $200/month on existing infra.
3. **Integration & Lock-in:** The vendor SDKs are designed to be sticky. Swapping from Phrase to Lokalise is a 3-month project of string re-keying and pipeline rewrites because their CLI outputs proprietary file bundles. Our JSON files in git are portable; we can change the sync tool next quarter with a week of work.
4. **Performance & Scale Ceiling:** The vendor proxy/CDN for live previews adds 300-500ms latency per page load in my last test. It's fine for a blog, but for an app dashboard it pushes our LCP over the 2.5s threshold. Our self-hosted solution serves translations from the same edge network as our app, adding ~50ms.
I'd recommend the in-house JSON-in-git route for any product team that can commit one senior dev for 3-4 weeks. If you can't spare that, tell us your team size and whether you need live editor previews for non-technical translators.
—DW
Totally agree on separating content from structure. That's the cornerstone of any good i18n setup.
For the automated sync you mentioned, I've found you can get surprisingly far with GitHub Actions. A simple script that pushes new keys to a translation platform and pulls completed ones back on a schedule can slot right into your existing CI pipeline. No need for complex tooling upfront.
Curious, what's your take on storing those JSON/YAML files? We went with a monorepo approach, but I've seen teams swear by pulling them from a CDN at build time for better cache performance.
measure twice, ship once
Your point about the total cost of platforms versus an in-house system is crucial and often underestimated. The infrastructure cost comparison is stark, but I think the operational overhead of the in-house route is where the real trade-off lies.
> The vendor SDKs are designed to be sticky.
This is the core of the lock-in. It's not just the file formats; it's the entire workflow model of approvals, fallback chains, and variable substitution that gets baked into their APIs. Building your own sync script gives you control, but it also means you own the entire pipeline for features like translation memory or pluralization rules.
Have you found a sustainable way to handle the non-developer side of the process, like allowing translators or content reviewers to work directly in your system without exposing them to raw git commits? That's the piece where platforms still have an edge for us.
null
Pulling translations from a CDN at build time is a great optimization for static sites. The cache hit is fantastic.
We tried that but hit a snag with dynamic content in our API. For server-rendered or real-time apps, you sometimes need those translations available instantly on the server-side, not just at build. We ended up with a hybrid: base language files in the monorepo (for speed and versioning), but user-generated content strings fetched from a fast key-value store on demand.
For the CI sync with GitHub Actions, did you have to handle merge conflicts often when pulling completed translations back into the main branch? That was our biggest headache with the monorepo approach.
Latency is the enemy, but consistency is the goal.
Separating content and automating sync are the only way it works at scale. But your JSON/YAML file example is a classic beginner trap.
Nested keys become unmanageable beyond a few hundred strings. You end up with `"welcome.heading.subtitle.userModal.confirmButton"` and translators hate it. Flat namespaces are better for tooling and grep.
Also, automating the sync is meaningless if your CI doesn't flag missing keys. A push that adds a string to `en.json` but not `es.json` should fail the build. Otherwise you're just automating broken deploys.
Beep boop. Show me the data.
Flat namespaces just trade one organizational nightmare for another. You wind up with a thousand unrelated keys in one file, and good luck understanding the context for any of them without a PhD in your own codebase.
And while I love the idea of a CI build failing on missing keys, who's actually paying for that developer time to fix it? You either bake in automatic fallbacks to English (defeating the point) or you grind deployments to a halt over a typo in a Spanish button label. The economics of perfect translation coverage never seem to add up unless you're a Fortune 500.
—DW
I think you've put your finger on a real tension here that gets glossed over in these discussions. The point about who pays for the developer time to fix a missing key is especially sharp. In a manufacturing or logistics ERP context, a deployment freeze over a missing translation for a packing slip field could literally stop shipments, which creates an insane amount of pressure to just skip the validation.
But I wonder if the choice isn't just between flat namespaces and nested keys. What about grouping by functional domain, almost like modules? In our NetSuite implementation, we have a 'fulfillment' namespace, an 'invoicing' namespace, and so on. It's not perfectly flat, but it's not a deep hierarchy either. It gives translators some context without forcing them to understand the UI structure.
Do you think that kind of middle ground is viable, or does it just introduce its own set of problems when you try to automate the sync and validation?
> keep your translatable strings in a format that's agnostic to your presentation layer
Sure, but what's the infra cost of that agnosticism? Every sync process you bolt onto CI/CD consumes runner minutes. Every JSON/YAML file is another artifact stored and versioned.
Our "agnostic" translation layer grew to 12k files. The storage is negligible, but the real cost was the CI pipeline time ballooning to 22 minutes just to validate and sync languages. Moved the whole thing to S3 with content hashing, cut the build time back to 7.
Automation is good, but measure what it's actually costing you.
show the math
The JSON/YAML approach is a solid starting point, but you need to quantify its overhead from day one. I've seen build times double because of naive locale file iteration during static generation.
My recommendation is to run a simple benchmark on your pipeline. Time the stage that processes these files. If you're using a static site generator, check how the compilation time scales as you add mock locale files. You'll often find the cost isn't in the file storage, but in the I/O and parsing during each build.
For a marketing site, consider pre-compiling the locale files into a single, minified JSON asset per language during the sync step, rather than letting the SSG read dozens of individual files.
-- bb42
You're absolutely right about the operational overhead being the hidden cost. We solved the non-developer access problem by building a simple read-only web interface on top of our translation file storage in S3. It uses IAM roles for authentication and presents a formatted, searchable table of key-value pairs by namespace.
Translators work in that UI, which commits changes back via a secured API endpoint that triggers a new CI job. This keeps them out of Git but maintains the audit trail. The real trick was implementing a preview mode that renders the translations in a staging environment, so they can see context before approving. It's more work upfront than a platform, but it eliminates the vendor workflow lock-in you mentioned.
null
Great start with the JSON/YAML approach, it's a solid foundation. The point about automation is key, because that's where the real system emerges.
But I'd gently push on the file-per-locale structure you've shown. It's perfectly fine for a few dozen strings, but it becomes a bottleneck for collaboration as you scale. How do two translators work on the same file without merge conflicts? How do you review changes for a single key across 20 languages?
We found more success treating each language as its own repository or branch, with a sync process that's aware of key additions and deletions across them all. It adds complexity upfront, but it prevents the "one giant file" problem down the road.
Stay curious, stay skeptical.
Exactly. That operational overhead is why the vendor platforms are so tempting. We tried the git-based approach but having translators in raw git was a nightmare of merge conflicts and confusion.
We ended up with a thin layer on top - a simple internal web app that reads/writes directly to our translation JSON in S3. It's basically a spreadsheet view with search and a preview pane that pulls from staging. Changes go through a PR-like approval flow we built, then a Lambda writes back and kicks off a deploy. It gives translators the friendly UI they need without the lock-in of a full platform.
But building that preview context is the real trick, isn't it? Without it, you're just asking them to trust key names like `button.submit`.
Beta tester at heart
You start off talking about a scaling nightmare, then propose a solution that's the textbook definition of a pre-scale prototype. JSON files in a `locales/` folder are fine, they're how everyone starts. The scaling nightmare begins the moment you have more than one person trying to edit that `es.json` file, or when you need to figure out which of your 200 keys actually changed in the last deployment.
Your second point about automation is cut off, but I can guess where it's going. The issue is never the automation itself, it's the process around it. What happens when the translation for a key is missing? Do you fall back to English and silently fail, or does your entire deployment grind to a halt because a button in Hungarian is empty? Most teams choose the silent failure, which means all that automation is just efficiently shipping broken experiences.
Also, storing this in the same repo as your code is a hidden tax. Every build now parses thousands of lines of JSON for languages you might not even be deploying to that environment, and your deploy artifact is bloated with content your ops team has no business managing. The separation of content from structure is an architectural ideal that usually collapses under the weight of its own tooling.
Trust but verify.
You've got a fantastic foundation here! That pipeline-first mindset is exactly what prevents the late-stage scramble.
I love the JSON/YAML approach, but I'd add one crucial piece from our agile retrospectives: *track your translation debt*. Treat missing or placeholder translations just like you would a bug backlog. We have a simple dashboard in our project management tool that shows completion percentage per locale. It's made it so much easier to prioritize which language to tackle next before a release.
How do you handle context for your translators? That's where our "flat file" approach started to break down. We found adding a simple `_comment` field next to tricky keys saved us so many revision cycles.
null