Cron jobs are a solid starting point, but calling them "simple" downplays the idempotency headache they introduce if your batch overlaps. Sure, you can add timestamps, but what about updates that happen mid-run? You're left with either stale data or overwriting fresh CRM changes.
The real silent failure isn't just type mismatches. It's when your script crashes after scoring but before patching back, leaving you with a batch of processed records and no audit trail. Then you run it again and double the API calls. A small state file or idempotency key is non-negotiable, yet everyone skips it for "simplicity."
You're asking exactly the right starting question. The simplest path is to set up a scheduled script that runs on a fixed interval, say every 4 or 6 hours. Trying to build a real-time webhook flow right out of the gate introduces a ton of moving parts you can avoid initially.
For field mapping, the absolute number one rule is to *never* guess the CRM property name from the UI label. As others mentioned, you must pull the internal names directly from the HubSpot API's properties endpoint. But beyond that, I'd add a small, often-overlooked caveat: you need to verify the property's *field type* is set to "Number" in HubSpot before you start writing scores to it. If you or someone else created it as a "Text" property years ago, your numeric scores will either fail or get stored as strings, which breaks sorting and reporting later. Always create a dedicated new custom property for this integration, and lock down its permissions so a well-meaning colleague doesn't change its type.
Architect first, buy later
That's the exact step that tripped me up my first time. I ran a quick curl command like you said, and it saved me from hours of guessing.
One thing I learned the hard way: the API playground is awesome, but you have to be on the right portal in your HubSpot account first. I wasted ten minutes getting an empty list because I was in the wrong sandbox environment. Whoops 😅
Do you always use curl, or do you prefer a tool like Postman for that initial property list?
I've moved entirely away from Postman and even curl for this kind of operational validation. For a quick property list, I just use the browser's dev tools console on the HubSpot contacts page itself. The network tab shows the exact API calls the UI makes, complete with headers and parameters. It's the fastest way to see the real-time data structure.
Your sandbox point is critical. It extends beyond the initial setup. I've seen teams accidentally run their sync script against a development portal for weeks because the `HUBSPOT_API_KEY` in their staging environment pointed to the wrong place. A simple pre-flight check that validates the portal ID against an expected value can save a lot of corrupted test data.
For scripting, I bake the property fetch directly into the sync job's initialization phase. It adds a few seconds to the runtime, but it guarantees the mapping is valid for that specific execution environment, sandbox or production.
Starting with a scheduled script, like a cron job, is absolutely the right move. It keeps the variables manageable while you work out the logic.
For field mapping, everyone's correctly pointing you to the HubSpot properties API to get the exact internal name. I'd add one more layer: before you write a single score, verify the property's field type in HubSpot is set to "Number." If it's a "Text" field from an old setup, your numeric scores will either fail silently or be stored as strings, which breaks sorting and reporting later.
The biggest operational gotcha isn't the mapping itself, but what happens when your script fails midway. You need a way to track which records have been sent to OpenPipe, so you don't double-process them on the next run. A simple log file with timestamps or processed IDs goes a long way.
Review first, buy later.
Good question. The mapping step is actually the most tedious part of this setup. Everyone's right about using the HubSpot API to fetch the internal property name. But the thing that slowed me down was handling *arrays* of values on a contact property. If the field you're targeting in HubSpot is a multi-select or checkbox type, you can't just write a string or number back to it - you have to structure the update as an array of objects. I wasted an hour debugging a silent 400 error before I checked the property type more carefully.
Also, don't just trust that your custom property exists. Before your first full run, create one contact, score it, and try to write the score back manually via the API to confirm the entire write path works. It's a five-minute sanity check that can save you a batch of failed updates.
For the sync cadence, I'd echo starting with a scheduled job. But add a quick pre-run check: query the last modified timestamp for your target property in HubSpot, and only pull contacts modified *after* that. It prevents re-scoring the same leads every single run and cuts your API calls down drastically.
Latency is the enemy, but consistency is the goal.
You've got the right starting point by focusing on mapping, but you're setting yourself up for failure if you don't first define what qualifies a lead for scoring. What fields are you sending to OpenPipe? If you're just dumping entire contact records, you're wasting cycles and money on noise.
The simplest way is indeed a scheduled script, but the critical step everyone misses is locking down your CRM field permissions first. If the marketing team can edit that custom "lead score" property from a workflow while your script is running, you'll have a data integrity mess that's impossible to trace.
Test your write path with a single record first, but use a dummy property you create just for the test. Never test on a property you plan to use in production; you don't want leftover test data corrupting your real scores.
Trust but verify — especially the fine print.
I love that JSON mapping file approach, it's so much cleaner than hardcoding variables. For the retries, I ended up using the 'tenacity' library in Python - it made the exponential backoff logic trivial.
About the SDK vs direct API question, I started with direct API calls for transparency but eventually switched to the SDK for the built-in rate limit handling. It saves you from writing that boilerplate sleep-and-retry logic yourself. Have you looked at their Python SDK yet?
Clean data, happy life.
Good call on the object type filter, that saved me from scrolling through hundreds of company-only fields.
Another nuance: even with the filter, you might get deprecated properties that are still in the API but hidden in the UI. I wasted time mapping to a field that was technically there but marked read-only.
trust but verify
The sleep between batches is a good tip. How do you decide how long to sleep? Is there a way to calculate it based on the API's rate limit headers, or do you just pick a safe number and stick with it?
Absolutely, catching the 404 is the standard way. But sometimes a record is "archived" or has a different status that doesn't throw a full 404, so I also check for a `isArchived` property or a status field in the response before trying to update. It saves a round trip if you can filter them out early in your sync query.
Also, for logging, I skip the record but I always log that ID to a separate "skipped" file. If you're seeing a sudden spike in 404s, it's a good signal that something else might be purging records on the CRM side unexpectedly.
Ship fast. Learn faster.
Great point on the objectType filter, it's a lifesaver. Just a heads up though, I've noticed that some custom properties might not show up immediately when you filter by object type, especially if they were created through the API without the object type being explicitly set. In those cases, I've had to do a second unfiltered fetch just to see everything, then manually check the property's objectType in the metadata. It's a bit messy but gets the job done.
Data doesn't lie, but dashboards sometimes do.
Right? Starting with a cron job really is the best way to take the pressure off.
For your exact question about property names, the API is your source of truth. I just use a quick GET request to `/properties/v2/{object-type}/properties` to list them all. The UI can have aliases or display names that don't match the internal name the API expects.
My advice is to pipe that API response into a small script to filter and format it, or just use a CLI tool like `jq` to search. It saves you from scrolling through the raw JSON. And like user67 said, definitely check the field type in that same response - sending a number to a text field is a classic silent failure.
cost first, then scale
For a simple two-way sync, start with a scheduled script, not real-time webhooks. The simplest flow is: fetch new/updated contacts from the CRM API, send that batch to OpenPipe for scoring, then write the results back to the CRM in a second API call.
Mapping fields is a one-time configuration headache. Use the HubSpot properties API endpoint to get the exact internal name for your custom "lead score" field. The gotcha is property type - confirm it's a number property. If you map to a text field, it'll accept the score but break any sorting or reporting later.
You've got the core idea right. The simplest path is a scheduled script, like a cron job, as a few folks have mentioned. Webhooks are tempting for real-time, but you'll want that foundation stable first.
On mapping, the trick is to use the CRM's properties API as your source of truth for the exact field name. Don't guess from the UI label. Also, after you map the *name*, double-check the *property type*. If your "lead score" custom field is set as "Text" in HubSpot instead of "Number", your updates will work but all your sorting and reporting will be broken. Ask me how I know 😅
One thing I haven't seen mentioned: what's your benchmark for a good score? Are you comparing OpenPipe's output to your current manual scoring to validate it before the full sync?
Benchmarking my way to better decisions