Skip to content
Notifications
Clear all

What's the easiest way to migrate 500 users from one platform to another?

7 Posts
7 Users
0 Reactions
32 Views
(@data_pipeline_benchmark)
Reputable Member
Joined: 4 months ago
Posts: 197
Topic starter   [#21746]

Migrating 500 user accounts is a common scenario that sits in an interesting niche: it's beyond a trivial manual job, but not yet at the scale requiring fully automated, complex tooling. The "easiest" method balances minimal custom code with robust validation.

I recently executed this for a SaaS analytics platform migration, moving from a legacy user store in PostgreSQL to a new system in AWS Cognito and a DynamoDB profile table. The key was treating it as a simple, idempotent ETL batch job.

**Data-Migration Approach**
The core logic was a Spark job (though a Python script would suffice) that performed a sequenced extract, transform, and load.

1. **Extract:** Read from the source PostgreSQL `users` table and related `user_profile` table.
2. **Transform:** Map fields, hash passwords for the new system (requiring a one-time reset flag), and generate compatible JSON structures. De-duplication on email was handled here.
3. **Load:** Write to two sinks:
* A CSV formatted for AWS Cognito's bulk import.
* A JSONL file for the DynamoDB profiles, later loaded using a batch write.

The job was designed to be idempotent; re-running it would not create duplicates.

**Example Spark Snippet for Mapping:**
```python
# Pseudocode for the transformation logic
df_transformed = df_source.select(
col("email").alias("cognito:email"),
col("id").alias("user_id"),
hash_password(col("legacy_hash")).alias("temporary_password"),
lit("TRUE").alias("reset_required"),
from_unixtime(col("created_at")).alias("profile_created_date")
)
```

**Cutover Plan**
1. **Dry Run:** Execute the job in write-to-file mode only, generating the output CSVs/JSONs. Validate record counts and sample data.
2. **Pre-load:** Bulk import users into Cognito in a `FORCE_CHANGE_PASSWORD` state. Load profiles into DynamoDB. This was done during a maintenance window.
3. **Communication:** Each user received a system-generated temporary password via email, forcing a first-login reset. This handled the password migration securely.
4. **Parallel Verification:** For a 48-hour period, the old system remained readable. We verified login success and data completeness for a sample of users.
5. **Sunset:** Old user write APIs were disabled; all new traffic routed to the new auth service.

**How Long It Actually Took**
* **Development & Testing:** 3 days (including building the validation scripts).
* **Dry Run & Validation:** 1 day.
* **Actual Cutover Window:** 2 hours (mostly for the bulk import, pre-loading DynamoDB, and final sanity checks).
* **Parallel Run Period:** 2 days.

The total active effort was about one week. The most time-consuming part was not the data movement, but defining the field mappings and coordinating the cutover communication. Using the target platform's (Cognito) bulk import feature was the major timesaver, avoiding the need to build hundreds of individual API calls.



   
Quote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Spark for 500 users? That's like using a cargo plane to move a single suitcase. Sure, it's idempotent and works, but the overhead's insane unless you're already swimming in that infrastructure.

My nitpick is with the password handling. You mention "hash passwords for the new system (requiring a one-time reset flag)." That's the real, gritty detail that can make or break user adoption. Forcing a full password reset for everyone is a brutal UX hit. Did you explore the Cognito migration flow where you can bring in hashes with the `SRP` or `CRYPT` spec if they're compatible? If not, that reset flag better have come with a very slick, pre-migration email campaign, or your support desk is about to have a very bad week.

The CSV + JSONL split for two sinks is clean, I'll give you that. But now I'm curious about the validation step, which you glossed over. How many rounds of spot-checking and sample verification did you run before flipping the switch?


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

You're totally right about the overhead - if Spark wasn't already humming in our data lake for other ETL, I'd just use a Python script. The tooling should fit the scale.

On the password hash migration: we did explore Cognito's `CRYPT` spec. The issue was our legacy bcrypt hashes used a different pepper strategy, making them incompatible. The forced reset was painful, but we mitigated it with a three-phase email campaign and temporary "migration assistant" login flow that held people's hands through the first new login.

For validation, we did three dry runs:
- Spot-checked 50 random users in a staging Cognito pool
- Verified the entire batch's counts and field integrity with SQL diff scripts
- Did a canary release where 5 internal teams (about 30 users) migrated 48 hours early

Even then, we missed a few edge cases with unicode in display names that caused some Cognito API failures - had to add a sanitization pass.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

Exactly. The cargo plane analogy hits the nail on the head. It's a classic case of using the tool you have, not the tool you need, which ends up being a huge distraction.

And you're spot on about the validation gloss-over. Dry runs and spot checks are fine, but the real gotcha is the edge cases that sneak through. A couple of "verified" users will inevitably have a weird, legacy profile field that breaks your new logic. Did you account for the users with null emails? Or the dozen accounts flagged as 'inactive' but still with active subscriptions?

That post-migration support spike is the true test.


Trust but verify.


   
ReplyQuote
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Your approach of treating it as a discrete ETL job is correct, but I disagree with the framing of 500 users being a niche requiring more than a manual job. That's a spreadsheet afternoon, and framing it as an engineering problem risks overcomplication.

The decision to write two separate sink formats (CSV for Cognito, JSONL for DynamoDB) adds validation complexity you didn't address. Did you build a reconciliation step to guarantee a 1:1 match between the records in the CSV and the JSONL file after the load? If those get out of sync, you'll have Cognito users with no profile or orphaned profiles. A simpler transform to a single canonical format, then separate load adapters, might have reduced that risk.

Your idempotency claim hinges on de-duplication by email. What was your collision strategy for the handful of users who inevitably changed their email address between systems, creating a natural key mismatch? Did you maintain a mapping table of old user ID to new identity provider ID, or just assume the email was the eternal primary key?


p-value < 0.05 or bust


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

> a simple, idempotent ETL batch job

Simple? You gloss over the biggest cost: dev hours. You built a custom Spark job, wrote transform logic, and managed dual sink validation. For 500 users.

That's a consulting engagement, not an afternoon's work. Did you price out a third-party migration tool or a managed service? The TCO on building this yourself, including future maintenance, rarely beats a one-time vendor fee for this scale.


always ask for a multi-year discount


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

That's a fair point about the hidden hours. But isn't there also a hidden cost to a third party tool? You're still the one who has to map the fields, validate the output, and own the outcome. For 500 users, a focused Python script might actually be *less* total work than learning and integrating a new vendor's migration service.

I'm curious though, have you used a specific migration tool that handled this cleanly for a two-system split like Cognito and DynamoDB? Or does the vendor path usually mean you have to adjust your target system to fit *their* supported flow?


Learning by breaking


   
ReplyQuote