Docs / Components / Streaming

Streaming

Real-time token streaming and incremental response processing with streamChat.

Langchain.Core.Model Langchain.Core.Stream

Real-time token streaming and incremental response processing with streamChat.

Key Concepts

  • Incremental Delivery: Stream tokens progressively as generated by the LLM for high-responsiveness and low perceived latency.
  • Stream Handlers: Pass a callback or consumer function to handle each chunk or token as it arrives over the wire.
  • Cancellation & Safety: Deterministic stream termination respecting Haskell’s exception and resource safety semantics.

Working Code Example

Compare local execution via Ollama and cloud API execution via OpenAI / OpenRouter. Use the toggle tabs or the global provider switcher in the header to switch:

{-# LANGUAGE LambdaCase #-}
{-# LANGUAGE OverloadedStrings #-}

module Ollama.Stream (runApp) where

import Control.Monad.Except (runExceptT)
import Control.Monad.IO.Class (liftIO)
import Control.Monad.Trans.Resource (runResourceT)
import Data.Conduit (runConduit, (.|))
import qualified Data.Conduit.List as CL
import qualified Data.Text.IO as T
import Langchain.Prelude
import System.IO (hFlush, stdout)

runApp :: IO ()
runApp = do
  o <- newOllama "gemma3" defaultConfig
  let msgs = [userMessage "What is the meaning of life"]
  res <- runResourceT $ runExceptT $ runConduit $ stream o msgs Nothing .| CL.mapM_ onStreamEvent
  case res of
    Left err -> T.putStrLn $ "\nError: " <> errorMessage err
    Right () -> T.putStrLn "\n--- Stream Finished ---"
  where
    onStreamEvent = \case
      LLMChunk _ chunk _ -> liftIO $ do
        T.putStr chunk
        hFlush stdout
      _ -> pure ()
{-# LANGUAGE LambdaCase #-}
{-# LANGUAGE OverloadedStrings #-}

module OpenAI.Stream (runApp) where

import Control.Monad.Except (runExceptT)
import Control.Monad.IO.Class (liftIO)
import Control.Monad.Trans.Resource (runResourceT)
import Data.Conduit (runConduit, (.|))
import qualified Data.Conduit.List as CL
import qualified Data.Text.IO as T
import Langchain.Prelude
import OpenAI.Common (defaultModelName, getOpenRouterModel)
import System.IO (hFlush, stdout)

runApp :: IO ()
runApp = do
  o <- getOpenRouterModel defaultModelName
  let msgs = [userMessage "What is the meaning of life"]
  res <- runResourceT $ runExceptT $ runConduit $ stream o msgs Nothing .| CL.mapM_ onStreamEvent
  case res of
    Left err -> T.putStrLn $ "\nError: " <> errorMessage err
    Right () -> T.putStrLn "\n--- Stream Finished ---"
  where
    onStreamEvent = \case
      LLMChunk _ chunk _ -> liftIO $ do
        T.putStr chunk
        hFlush stdout
      _ -> pure ()

Core Types & Functions

ChatModel m => m -> [Message] -> (Text -> IO ()) -> ExceptT LangchainError IO Message

Running This Example

Local Ollama

Ensure your Ollama daemon is running locally with the target model:

ollama run gemma3 # or your desired model
stack run streamollama

OpenAI / OpenRouter

Ensure your OPENROUTER_API_KEY or OPENAI_API_KEY is exported:

export OPENROUTER_API_KEY="your-api-key"
stack run streamopenai
ESC