I'm working on a Python Flask API at work and need to generate OpenAPI specs. Doing it manually is taking forever.
Has anyone used Claude Code for this? I'm curious if it can reliably read my route definitions and decorators to build the spec. Did it handle complex request/response models well? Any pitfalls I should watch for?
I tried Claude Code for something similar with FastAPI last month. It got the basic structure right but missed some Pydantic field validators in my nested models.
One pitfall: it sometimes hallucinates example data that doesn't match my actual schema. I ended up having to manually check every generated example.
Have you looked at using a library like flask-smorest or apiflask instead? They build the spec automatically from your decorators. Might save you more time than generating from scratch.
Containers are magic, but I want to know how the magic works.
I ran a similar test for documenting our internal HR API endpoints, also a Flask setup. Claude Code did a decent job pulling basic route parameters and response codes from my decorators.
But I noticed it struggled with conditional logic in the specs, like different required fields for different user roles. It would flatten everything into one schema. Did you encounter anything like that in your testing?
I ran into the exact same issue with the hallucinated examples. They seem plausible until you realize the enum values are wrong or a required nested object is missing.
Your library suggestion is the right call. I've found that even when the AI gets 80% of it right, that last 20% of manual validation and correction kills the ROI. A purpose-built library like flask-smorest might need initial setup, but you're getting a guaranteed correct spec and it stays in sync automatically.
You end up spending more time prompting and fixing than just using the right tool from the start.