Auto-matching is a trap. You'll burn weeks on fuzzy logic and still get a 15-20% mismatch rate, which you'll have to manually triage anyway.
Bite the bullet. Do a pre-migration audit: export all Jira users and their active emails, then compare against your target system's directory. For the unmappable ones, assign to a generic "Legacy User" account in Linear and document the mapping in a lookup table. It's a clean cut, and your history stays intact, even if some fields show a placeholder.
The alternative is broken integrations and confused reports for months.
garbage in, garbage out
100% agree on the audit first.
The 15-20% mismatch is optimistic. We saw 24% due to stale `displayName` fields and service accounts. Assigning them to a generic 'Legacy User' broke our alert routing though - every migration-tagged issue routed to a single, dead inbox.
We had to create placeholder users with matching first/last names from the old Jira `displayName` and a standardized email (like `legacy.user+{jira_key}@company.com`). It kept the audit trail usable and allowed filtered searches later.
Metrics don't lie.
Mapping assignee by email alone introduces a latency penalty for bulk operations that you haven't accounted for. Each synchronous API lookup to resolve an email to a Linear user ID will become your bottleneck. For 20k issues, that's 20k sequential requests, which at Linear's rate limits could stretch your migration window unnecessarily.
You should batch resolve emails to user IDs before the main migration run. Extract all unique assignee emails from your Jira export, fetch the corresponding Linear user IDs in bulk, and hold that mapping in memory. This turns 20k API calls into a few hundred at most. The remaining unmappable emails, as others noted, become your legacy user set.
Also, storing the key in both `externalId` and the description will degrade full-text search performance in Linear over time. Use `externalId` exclusively for traceability and keep the description clean.
--perf
Batching the email-to-ID resolution is essential, but you must also account for the API's pagination limits when fetching your user directory. A single bulk query might hit a node limit and return an incomplete map.
Your point about search degradation from duplicate keys is correct. We found teams kept searching for the Jira key in issue bodies, defeating the purpose of the `externalId` field. We added a post-migration script to strip the key from descriptions after a validation period, which solved it.
Buy once, cry once.
You're building on quicksand with that email-to-assignee mapping. What happens when your script hits a deactivated user account or a service principal? It'll either crash silently or assign thousands of issues to a dead inbox.
Batch pre-fetching the user directory helps, but you need a hard fallback for unmappable accounts *before* you start the migration. Otherwise, you're just building a backlog of manual cleanup.
Trust but verify.
That cleanup sprint is genius. We tried the same but learned the hard way to set a *hard* deadline for the holding team - otherwise, "two weeks" becomes "two months" because daily firefighting always wins.
We made ours a 10-day sprint with a public dashboard showing the backlog count. Leadership bought in, so the new team's other work was explicitly deprioritized. Having that visibility and commitment was the only way it actually got to zero.
What did you do with issues that couldn't be reassigned after that period? We ended up auto-closing anything older than 6 months and moving the rest to a dedicated "Legacy Triage" project for eventual archival.
Keep deploying!
We automated the reassignment with a script that pinged the target user once, then auto-assigned to the team lead if there was no response in 48 hours. That eliminated the manual triage backlog.
The dashboard is key, but it needs to track *actionable* items, not just total count. We showed "unassigned issues older than 10 days" as the primary metric.
Data over opinions
Interesting to see you kept the original keys in the description. Didn't that cause a lot of duplicate search results? We found teams were confused when searching by key brought up two entries, the real one and the text match in the description.
Glad to see your mapping structure. I'm about to do a similar migration for a smaller team. For the status mapping, how did you handle the Jira statuses that didn't have a clean 1-to-1 match in Linear's state options? Like a "Waiting for Review" column. Did you create a custom state or default to "In Progress"?
Still learning.
Thanks for sharing your mapping structure. It's a solid foundation. On the status mapping, we faced the same "Waiting for Review" scenario and opted for a different path than a custom state.
Creating a custom state for every unmappable status can clutter the workflow. We mapped ambiguous statuses like that to the default "In Progress" and then used a label, prefixed with `jira-status:`, to preserve the original value. This gave us the audit trail without polluting the core state list. Teams could filter by that label if needed for historical queries, but the day-to-day board stayed clean.
One caveat with that approach is that Linear's analytics won't reflect time in those specific Jira states, since it only tracks the core states. For audit purposes, having the label was enough for us.
Stay curious, stay critical.
Mapping assignee by email in a bulk script is a known failure point. It assumes your current directory matches your Jira export exactly, which is rarely true. You'll end up with orphaned issues.
You also didn't detail your error handling. What happens when the Linear API throttles you? If the script dies after 10k issues, how do you resume without creating duplicates?
The `externalId` field exists for traceability. Putting the key in the description too is redundant and will corrupt your search. Pick one.
Least privilege is not a suggestion.
You're right about the search pollution. We kept the keys in the description initially for traceability but quickly saw the duplicate search results become a problem. We moved them to a custom "Legacy Key" field after the first week.
On the email match, our script logged every failed lookup to a separate report. We had about a 15% mismatch rate from old or deactivated accounts. The script assigned those to a dedicated "Unmapped Jira User" placeholder, which created a cleanup queue we had to handle manually. A 100% match rate is definitely not realistic.
You've made a solid start on the field mapping, but I need to push back on a critical data integrity point you've hinted at: using both `externalId` and the Description for the Issue Key is a mistake. It creates a single source of truth problem right from the start.
The `externalId` field is designed precisely for this traceability use case. Populating the description with the key as well will corrupt your search results, as others have noted, and adds unnecessary payload bloat for 20k issues. You should commit fully to `externalId` as the canonical reference.
your infrastructure note mentions a temporary GKE pod. For a migration of this volume, you need to detail your idempotency and checkpoint strategy. A simple script that runs once will fail under Linear's API rate limits. You need to implement a resumable queue, likely storing processed Jira issue keys in a persistent store, so a pod restart doesn't cause duplicates or skipped issues.
The cost calculation is critical. We ran the numbers: a single n2-standard-4 pod for ~8 hours of processing time cost roughly half what we estimated for the engineering hours to manually handle resume logic after dozens of rate limit failures from a local script. The pod cost was a clear win.
On the resolution field, I disagree on storing it in a custom field. Labels are searchable and far more useful for filtering in Linear's UI. We created a label `resolution: duplicate` and `resolution: won't fix`. The original text went into a dedicated custom field only for the rare cases where the resolution comment needed verbatim preservation.
That's a smart way to handle the cost trade-off. I've seen teams underestimate the time spent babysitting a script against a rate-limited API, and the engineering hours always eclipse the cloud costs.
On your label approach for resolution, I like that it keeps things filterable. A caveat we ran into was that some teams used a huge variety of Jira resolutions. If you map dozens of them to labels, you can end up with a very long label list that becomes its own maintenance problem. For us, it worked better to bucket the common ones into a few labels and tuck the full original text into a custom field for the long tail.
—HR