Streaming
Real-time token streaming and incremental response processing with streamChat.
Real-time token streaming and incremental response processing with streamChat.
Key Concepts
- Incremental Delivery: Stream tokens progressively as generated by the LLM for high-responsiveness and low perceived latency.
- Stream Handlers: Pass a callback or consumer function to handle each chunk or token as it arrives over the wire.
- Cancellation & Safety: Deterministic stream termination respecting Haskellβs exception and resource safety semantics.
Working Code Example
Compare local execution via Ollama and cloud API execution via OpenAI / OpenRouter. Use the toggle tabs or the global provider switcher in the header to switch:
{-# LANGUAGE LambdaCase #-}
{-# LANGUAGE OverloadedStrings #-}
module Ollama.Stream (runApp) where
import Control.Monad.Except (runExceptT)
import Control.Monad.IO.Class (liftIO)
import Control.Monad.Trans.Resource (runResourceT)
import Data.Conduit (runConduit, (.|))
import qualified Data.Conduit.List as CL
import qualified Data.Text.IO as T
import Langchain.Prelude
import System.IO (hFlush, stdout)
runApp :: IO ()
runApp = do
o <- newOllama "gemma3" defaultConfig
let msgs = [userMessage "What is the meaning of life"]
res <- runResourceT $ runExceptT $ runConduit $ stream o msgs Nothing .| CL.mapM_ onStreamEvent
case res of
Left err -> T.putStrLn $ "\nError: " <> errorMessage err
Right () -> T.putStrLn "\n--- Stream Finished ---"
where
onStreamEvent = \case
LLMChunk _ chunk _ -> liftIO $ do
T.putStr chunk
hFlush stdout
_ -> pure (){-# LANGUAGE LambdaCase #-}
{-# LANGUAGE OverloadedStrings #-}
module OpenAI.Stream (runApp) where
import Control.Monad.Except (runExceptT)
import Control.Monad.IO.Class (liftIO)
import Control.Monad.Trans.Resource (runResourceT)
import Data.Conduit (runConduit, (.|))
import qualified Data.Conduit.List as CL
import qualified Data.Text.IO as T
import Langchain.Prelude
import OpenAI.Common (defaultModelName, getOpenRouterModel)
import System.IO (hFlush, stdout)
runApp :: IO ()
runApp = do
o <- getOpenRouterModel defaultModelName
let msgs = [userMessage "What is the meaning of life"]
res <- runResourceT $ runExceptT $ runConduit $ stream o msgs Nothing .| CL.mapM_ onStreamEvent
case res of
Left err -> T.putStrLn $ "\nError: " <> errorMessage err
Right () -> T.putStrLn "\n--- Stream Finished ---"
where
onStreamEvent = \case
LLMChunk _ chunk _ -> liftIO $ do
T.putStr chunk
hFlush stdout
_ -> pure ()Core Types & Functions
ChatModel m => m -> [Message] -> (Text -> IO ()) -> ExceptT LangchainError IO MessageRunning This Example
Local Ollama
Ensure your Ollama daemon is running locally with the target model:
ollama run gemma3 # or your desired model
stack run streamollamaOpenAI / OpenRouter
Ensure your OPENROUTER_API_KEY or OPENAI_API_KEY is exported:
export OPENROUTER_API_KEY="your-api-key"
stack run streamopenai