Skip to content
Notifications
Clear all

Step-by-step: Onboarding a legacy mainframe system

54 Posts
52 Users
0 Reactions
154 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

>learning 3270 and RACF on the fly

This was our exact situation. Zero institutional mainframe knowledge. We started with the product manuals for the 3270 emulator we were automating, treating it as just another headless protocol to script. The bigger hurdle was getting the system team to grant the PAM service ID the correct, minimal RACF permissions it needed to start a session, which required a negotiation we weren't prepared for.


benchmark or bust


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

>getting the system team to grant the PAM service ID the correct, minimal RACF permissions

This is the negotiation that defines the project timeline. The mainframe team's default stance is often "no new IDs, ever" because their uptime metric is based on stability, not feature velocity.

We had to build a proposal showing the exact RACF commands, tying each one to a specific, auditable screen in our automation flow. It's like bringing a detailed spec to a change control board for a system they see as finished. The goal isn't just access, it's proving your request is the most boring, risk-free change they'll see all quarter.



   
ReplyQuote
(@diego_h)
Honorable Member
Joined: 6 months ago
Posts: 313
 

That latency cost hits hard. We're building a connector for order status checks, and the added delay pushes our real-time promise out the window. It's not just batch windows.

Did you find any workaround for the scraping delay itself? We're thinking of caching screen layouts, but if the screen logic changes based on data, that seems risky.


Still learning.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Caching screen layouts helped us shave off a second or two, but you're right to be nervous about logic changes. We had to get clever and cache *conditionally*.

Our scraper first checks a few key, static screen attributes (like the screen ID in the upper right) against our cache. If there's a match, it uses the cached field positions for the initial data grab. But then it runs a quick, secondary validation on the retrieved data against known patterns. If anything looks off, it defaults back to a full, fresh scrape for that session.

It added some complexity, but the conditional logic kept our response time tolerable for most routine queries. The real killer for us was the unpredictable screen transitions, not the layout itself.



   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

You've pinpointed the exact failure mode. Validating the data pipeline end-to-end is non-negotiable, but it's a step that's often skipped in the rush to declare a proof-of-concept successful.

We learned this the hard way when our custom session ID, while present in the proxy's JSON-structured logs, was being flattened into key-value pairs by a legacy log shipper. The field name contained a dot, which the shipper interpreted as a nested object delimiter and silently dropped. The auditors, working from the parsed data in Snowflake, saw nothing.

The lesson was to treat the log format specification as a critical integration contract, not an implementation detail. You need to test the stitch from the raw byte in your application all the way to the rendered column in the auditor's BI tool.


Measure twice, cut once.


   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

>budget triple the time and a lot of whiskey.

That's the starting assumption, not the ceiling. The mapping is one battle, but the operational reality is another. You built the connector, but now you have to monitor and maintain it. Suddenly you're running a 3270 client as a production service, with all its quirks.

Your alerts now need to distinguish between a mainframe outage, a network partition to the data center, a session timeout in the emulator, and the off chance someone actually changed a green screen layout. Each has a different escalation path, and none of them are in your runbooks. The whiskey is for when you get paged at 2 AM because the PAM service locked out a service ID after three failed scrapes, and the mainframe team won't even look at it until their 6 AM batch window closes.


Automate everything. Twice.


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

Completely agree on the separation of concerns. That split was the only way our project got past the security review. We used HashiCorp Vault for the credential lifecycle and a separate, home-grown session logger that just captured the raw 3270 stream with timestamps and the ephemeral user ID.

The critical part, and this took us months to lock down, was guaranteeing the integrity of the link between the two systems. The recorder had to receive a cryptographically signed token from the PAM system with the exact, one-time-use mainframe credentials and the internal request ID. Otherwise, you're just correlating logs based on timestamps, which an auditor will tear apart. The recorder's own logs of receiving that token became a control point.

If that handoff isn't atomic, you're left with a gap where someone could have accessed a session, but the recorder failed to start, leaving no evidence the key was ever used.


Logs don't lie.


   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

You're right, the mapping is a multi-dimensional puzzle. We had a similar struggle translating IMS TM transaction codes into modern role-based access. The connector can handle the API translation, but the real logic lies in reconciling the mainframe's procedural security model - where access is often tied to the *program* being run - with Delinea's identity-centric model. We ended up creating a middleware layer that acted as a policy decision point, interpreting the modern user's role to select the appropriate, granular mainframe transaction code for the session.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

That "triple the time" estimate is optimistic if you haven't factored in the ongoing compute burn for emulating a user at scale. Everyone gets hyper-focused on the mapping logic and forgets to cost out what it means to run a headless 3270 client in a loop.

You're essentially paying for a virtual employee to stare at a green screen and type. If your connector is polling for status checks every few minutes, you're provisioning a full-time instance for a task that takes milliseconds of CPU. I've seen teams spin up dedicated t3.mediums just to host a single-threaded Perl script that sleeps 95% of the time. The math on that, over a year, buys a lot of whiskey.


pay for what you use, not what you reserve


   
ReplyQuote
Page 4 / 4