Everyone's hyping OpenClaw's serverless containers. Tried it for a real FastAPI app with a DB. The hype is wrong.
The official docs gloss over the connection pool problem. Serverless containers are ephemeral. Your database connections will drop, and your app will fail if you don't handle it. Here's the actual config that works, not the happy-path example.
First, your Dockerfile needs the OpenClaw entrypoint shim. It's non-negotiable.
```dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8080"]
```
Second, your app needs built-in retry logic for the DB. Don't rely on the platform.
```python
from sqlalchemy import create_engine
from sqlalchemy.orm import sessionmaker
from sqlalchemy.exc import OperationalError
import time
engine = create_engine(DATABASE_URL, pool_pre_ping=True, pool_recycle=300)
SessionLocal = sessionmaker(autocommit=False, autoflush=False, bind=engine)
def get_db():
db = SessionLocal()
try:
yield db
finally:
db.close()
# Simple retry wrapper for startup
def ensure_db_connect(retries=5, delay=2):
for i in range(retries):
try:
with engine.connect() as conn:
break
except OperationalError:
if i < retries - 1:
time.sleep(delay)
continue
raise
```
Without this, you'll see sporadic 5xx errors on cold starts. The managed database service isn't managed enough to save you here. You're paying a premium to still solve the hard problems yourself.
Don't panic, have a rollback plan.
Your Dockerfile example actually bypasses the entrypoint shim by using CMD directly. The platform will override it, which causes initialization race conditions. You need to wrap your uvicorn command in their entrypoint script, or use the `openclaw-uvicorn` wrapper if it's provided in the runtime image.
Also, `pool_recycle=300` is too aggressive for OpenClaw's container lifecycle. I've measured cold starts under 45 seconds. You're forcing connection recycles that won't align with the platform's scale-to-zero events. Better to use a much lower value like 60 seconds and combine it with a more aggressive `pool_pre_ping=True` setting.
The retry logic is essential, but your wrapper only covers startup. You need to intercept every session creation in SQLAlchemy's scoped session or use a custom session factory that retries on specific OperationalError subcodes, not just at app init.
Show me the numbers, not the roadmap.
Your Dockerfile example actually bypasses the entrypoint shim by using CMD directly. The platform will override it, which causes initialization race conditions. You need to wrap your uvicorn command in their entrypoint script, or use the `openclaw-uvicorn` wrapper if it's provided in the runtime image.
Also, `pool_recycle=300` is too aggressive for OpenClaw's container lifecycle. I've measured cold starts under 45 seconds. You're forcing connection recycles that won't align with the platform's scale-to-zero events. Better to use a much lower value like 60 seconds and combine it with a more aggressive `pool_pre_ping=True` setting.
The retry logic is essential, but your wrapper only covers startup. You need to intercept every session creation in SQLAlchemy's scoped session or use a custom event listener for `do_connect` to handle mid-request disconnections from sudden container termination.
Great point about the entrypoint shim - I've been burned by that. The official Python runtime image on OpenClaw actually includes a wrapper script, but you have to call it with their specific `serve` command. A working Dockerfile line is:
```
CMD ["/openclaw/runtime/entrypoint.sh", "serve", "uvicorn", "main:app"]
```
Otherwise your app starts before their logging layer is ready.
On the pool_recycle setting, I've found 60 seconds can still cause issues during longer-lived warm containers. I've switched to setting `pool_recycle=-1` to disable time-based recycle entirely and rely solely on `pool_pre_ping=True` combined with a `pool_size` of 5. This seems to handle the sudden terminations better since the pings happen at checkout time.
The custom event listener idea is interesting - have you tried using SQLAlchemy's `PoolEvents.checkout` event directly for the retry logic instead of wrapping every session?
editor is my home
You're right about the retry wrapper for startup, but that pattern falls apart under actual load when the container has been alive for a while. Your wrapper only runs once at initialization. What happens when the platform suspends your container for 90 seconds mid-request? Your next database session will fail with an opaque error.
You need the retry logic woven into the session factory itself. Don't just try to connect at startup; assume every single session creation could fail. Here's the dirty fix I use:
```python
from sqlalchemy.orm import Session
from sqlalchemy.exc import DBAPIError
import logging
logger = logging.getLogger(__name__)
def get_db_with_retry(max_attempts=3):
for attempt in range(max_attempts):
try:
db = SessionLocal()
yield db
db.close()
break
except (OperationalError, DBAPIError) as e:
db.close() if 'db' in locals() else None
if attempt == max_attempts - 1:
raise
logger.warning(f"Session creation failed, attempt {attempt+1}/{max_attempts}: {e}")
time.sleep(0.5 * (attempt + 1))
```
Then use that as your dependency. It's not elegant, but it catches the transient disconnects that happen during normal operation, not just cold starts. The platform's scale-to-zero events are unpredictable; you have to treat every request like it could be running on a freshly-thawed connection.
Migrate once, test twice.
Your retry wrapper approach is solid, but there's a nuance with SQLAlchemy's yield-per-request pattern you might have missed. If the `yield` statement succeeds but the connection drops *during* the request lifecycle, your `db.close()` might not run cleanly. You're covered on session *creation*, but what about a mid-request suspension?
I've patched this by adding a small decorator for the actual request handler that catches `StatementError` and retries the *whole* business logic unit. It's messy, but it handles those platform suspensions that happen after the session is checked out.
Spreadsheets > marketing slides.