← back

Async & Concurrency for AI Engineers

Every real service — including one built around an LLM — eventually needs to do more than one thing at once: wait on several API calls, stream a response back token by token, or run slow work in the background without blocking a user. This is a tour through the tools Python actually gives you for that, and which one fits which job.

async & await Fundamentals

async def marks a function as awaitable — await is where it actually pauses and hands control back. Calling an async function without await just creates a coroutine object; it never runs until something awaits it. async/await doesn't run code in parallel — it lets one function pause and resume without blocking everything else.

import asyncio

async def fetch_price():
    await asyncio.sleep(1)
    return 0.5

async def main():
    price = await fetch_price()
    print(price)   # 0.5

asyncio.run(main())

The asyncio Event Loop & Cooperative Concurrency

The event loop runs one task at a time, switching between them only at each await. asyncio.gather runs coroutines concurrently — two 1-second waits finish in about 1 second total, not 2, because each await hands control back to the loop while waiting. asyncio gives you concurrency, not parallelism.

async def fetch_price(name, delay):
    await asyncio.sleep(delay)
    return name, delay

async def main():
    results = await asyncio.gather(fetch_price("A", 1), fetch_price("B", 1))
    print(results)   # takes ~1 second total, not 2

Async HTTP & Concurrent LLM/API Requests

Async HTTP lets you fire off many requests at once and collect results as each one finishes. Each network call spends almost all its time waiting, not computing — exactly the situation async concurrency is built for. A regular (non-async) HTTP library inside async code still blocks the event loop, since it doesn't know how to await.

async def call_llm():
    await asyncio.sleep(1)
    return "a ripe, sweet apple"

async def call_translate():
    await asyncio.sleep(1)
    return "una manzana"

description, translation = await asyncio.gather(call_llm(), call_translate())

Threads vs Processes vs Async

Threads share memory and switch anytime; processes run fully separately; async switches only at await. Threads can be interrupted at any instruction, mid-line even — async only switches at an explicit await, which is what makes async code easier to reason about. Async/threads for I/O-bound waiting; processes for CPU-bound crunching that needs real parallelism.

The Python GIL

The GIL lets only one thread run Python bytecode at a time, even on a multi-core machine. Adding more threads to a CPU-bound Python function usually makes it no faster, because the GIL still only lets one run at once. I/O waits (network, disk, sleep) release the GIL, so threads still help for I/O-bound work — just not for CPU-bound work.

io_bound_task = "waiting on 5 network calls"
answer = "no — I/O releases the GIL, threads still help"

cpu_bound_task = "crunching numbers in a loop"
answer = "yes — the GIL serializes it, threads won't help"

Background Jobs, Queues & Worker Patterns

A background job runs outside the request/response cycle — a queue is what hands it off. The web request only needs to enqueue the job and return; the actual work happens later, in a separate worker process, completely decoupled from the request.

import queue

job_queue = queue.Queue()

def handle_checkout(order):
    job_queue.put(("send_email", order))
    return "checkout complete"   # returns instantly

def worker():
    while not job_queue.empty():
        job = job_queue.get()
        print("processing:", job)

Streaming AI Responses

Streaming sends a response token by token as it's generated, instead of waiting for the whole thing. A streaming response is really just an iterator — each chunk arrives as its own piece, printed or sent the moment it's available. Buffering the entire response first makes a user stare at a blank screen for as long as the slowest full generation takes.

def stream_words(words):
    for word in words:
        time.sleep(0.2)
        yield word

for word in stream_words(["a", "ripe", "sweet", "apple"]):
    print(word, end=" ")

Choosing Concurrency: I/O-bound vs CPU-bound

The first question is always: is this work waiting, or computing? I/O-bound work — network, disk, sleep — spends its time waiting, so async or threads let those waits overlap. CPU-bound work keeps the CPU genuinely busy, so only separate processes actually run it faster.

task_a = "fetching prices from 10 suppliers over the network"
kind_a = "I/O-bound -> async or threads"

task_b = "scoring 10,000 items with a local model"
kind_b = "CPU-bound -> multiprocessing"

These pieces explain how a real service — including one built around an LLM — actually handles doing more than one thing at once. The fastest way to make them stick is to find one slow, blocking part of a real project and give it the right kind of concurrency.