Async & Concurrency for AI Engineers
Every real service — including one built around an LLM — eventually needs to do more than one thing at once: wait on several API calls, stream a response back token by token, or run slow work in the background without blocking a user. This is a tour through the tools Python actually gives you for that, and which one fits which job.
async & await Fundamentals
async def marks a function as awaitable — await is where it actually pauses and hands control back. Calling an async function without await just creates a coroutine object; it never runs until something awaits it. async/await doesn't run code in parallel — it lets one function pause and resume without blocking everything else.
import asyncio
async def fetch_price():
await asyncio.sleep(1)
return 0.5
async def main():
price = await fetch_price()
print(price) # 0.5
asyncio.run(main())The asyncio Event Loop & Cooperative Concurrency
The event loop runs one task at a time, switching between them only at each await. asyncio.gather runs coroutines concurrently — two 1-second waits finish in about 1 second total, not 2, because each await hands control back to the loop while waiting. asyncio gives you concurrency, not parallelism.
async def fetch_price(name, delay):
await asyncio.sleep(delay)
return name, delay
async def main():
results = await asyncio.gather(fetch_price("A", 1), fetch_price("B", 1))
print(results) # takes ~1 second total, not 2Async HTTP & Concurrent LLM/API Requests
Async HTTP lets you fire off many requests at once and collect results as each one finishes. Each network call spends almost all its time waiting, not computing — exactly the situation async concurrency is built for. A regular (non-async) HTTP library inside async code still blocks the event loop, since it doesn't know how to await.
async def call_llm():
await asyncio.sleep(1)
return "a ripe, sweet apple"
async def call_translate():
await asyncio.sleep(1)
return "una manzana"
description, translation = await asyncio.gather(call_llm(), call_translate())Threads vs Processes vs Async
Threads share memory and switch anytime; processes run fully separately; async switches only at await. Threads can be interrupted at any instruction, mid-line even — async only switches at an explicit await, which is what makes async code easier to reason about. Async/threads for I/O-bound waiting; processes for CPU-bound crunching that needs real parallelism.
- →Fetching from 5 suppliers over the network: async (or threads) — I/O-bound, mostly waiting
- →Recalculating a large in-memory pricing table: multiprocessing — CPU-bound, needs real parallelism
The Python GIL
The GIL lets only one thread run Python bytecode at a time, even on a multi-core machine. Adding more threads to a CPU-bound Python function usually makes it no faster, because the GIL still only lets one run at once. I/O waits (network, disk, sleep) release the GIL, so threads still help for I/O-bound work — just not for CPU-bound work.
io_bound_task = "waiting on 5 network calls"
answer = "no — I/O releases the GIL, threads still help"
cpu_bound_task = "crunching numbers in a loop"
answer = "yes — the GIL serializes it, threads won't help"Background Jobs, Queues & Worker Patterns
A background job runs outside the request/response cycle — a queue is what hands it off. The web request only needs to enqueue the job and return; the actual work happens later, in a separate worker process, completely decoupled from the request.
import queue
job_queue = queue.Queue()
def handle_checkout(order):
job_queue.put(("send_email", order))
return "checkout complete" # returns instantly
def worker():
while not job_queue.empty():
job = job_queue.get()
print("processing:", job)Streaming AI Responses
Streaming sends a response token by token as it's generated, instead of waiting for the whole thing. A streaming response is really just an iterator — each chunk arrives as its own piece, printed or sent the moment it's available. Buffering the entire response first makes a user stare at a blank screen for as long as the slowest full generation takes.
def stream_words(words):
for word in words:
time.sleep(0.2)
yield word
for word in stream_words(["a", "ripe", "sweet", "apple"]):
print(word, end=" ")Choosing Concurrency: I/O-bound vs CPU-bound
The first question is always: is this work waiting, or computing? I/O-bound work — network, disk, sleep — spends its time waiting, so async or threads let those waits overlap. CPU-bound work keeps the CPU genuinely busy, so only separate processes actually run it faster.
task_a = "fetching prices from 10 suppliers over the network"
kind_a = "I/O-bound -> async or threads"
task_b = "scoring 10,000 items with a local model"
kind_b = "CPU-bound -> multiprocessing"These pieces explain how a real service — including one built around an LLM — actually handles doing more than one thing at once. The fastest way to make them stick is to find one slow, blocking part of a real project and give it the right kind of concurrency.