DO: Iterate ON the Space. Push early, verify against live URL. Stream logs. Read logs first, act once, make one targeted fix. Use cheapest iteration rung. Set all three cache env vars at module top, before any import. Measure @spaces.GPU(duration=) empirically. Call c.view_api() before any predict() call. python3 -m py_compile app.py is the maximum local check before pushing. Always ship 3–5 curated gr.Examples and cache them (cache_examples=True, cache_mode="lazy", with fn=/outputs= set) — find good inputs, not placeholders.
NEVER: Sleep-poll. Build mock modes, SKIP_MODEL_LOAD env vars, Playwright harnesses, or local Gradio servers. Restart before reading the error. Issue concurrent uploads. Restart while uploading. Stack restarts while runtime.sha is still flipping. Use integer device IDs (.to(0), device_map={"": 0}, set_device(0)). Pin torch or torchaudio. Proxy another community ZeroGPU Space via `gradio_c