- Tool calling - calling external tools (like databases queries or API calls) and use results in their responses.
- Structured output - where the model’s response is constrained to follow a defined format.
- Multimodality - process and return data other than text, such as images, audio, and video.
- Reasoning - models perform multi-step reasoning to arrive at a conclusion.
Basic usage
Models can be utilized in two ways:- With agents - Models can be dynamically specified when creating an agent.
- Standalone - Models can be called directly (outside of the agent loop) for tasks like text generation, classification, or extraction without the need for an agent framework.
Initialize a model
The easiest way to get started with a standalone model in LangChain is to useinit_chat_model to initialize one from a chat model provider of your choice (examples below):
- OpenAI
- Anthropic
- Azure
- Google Gemini
- AWS Bedrock
- HuggingFace
- OpenRouter
init_chat_model for more detail, including information on how to pass model parameters.
Supported providers and models
LangChain supports all major model providers through dedicated integration packages. Each provider package implements the same standard interface, so you can swap providers without rewriting application logic. New model names work immediately — no LangChain update required — because provider packages pass model names directly to the provider’s API. Browse the full list of supported providers, or see Providers and models for a conceptual overview of how providers, packages, and model names work together in LangChain.Key methods
Invoke
Stream
Batch
Parameters
A chat model takes parameters that can be used to configure its behavior. The full set of supported parameters varies by model and provider, but standard ones include:init_chat_model, pass these parameters as inline :
Connection resilience
LangChain chat models automatically retry failed API requests with exponential backoff. By default, models retry up to 6 times for network errors, rate limits (429), and server errors (5xx). Client errors like 401 (unauthorized) or 404 are not retried. You can adjustmax_retries and timeout when creating a model, then pass that instance to create_agent, create_deep_agent, or call it standalone:
ChatOpenAI has use_responses_api to dictate whether to use the OpenAI Responses or Completions API.To find all the parameters supported by a given chat model, head to the chat model integrations page.Invocation
A chat model must be invoked to generate an output. There are three primary invocation methods, each suited to different use cases.Invoke
The most straightforward way to call a model is to useinvoke() with a single message or a list of messages.
ChatOpenAI(/oss/integrations/chat/openai).Stream
Most models can stream their output content while it is being generated. By displaying output progressively, streaming significantly improves user experience, particularly for longer responses. Callingstream() returns an that yields output chunks as they are produced. You can use a loop to process each chunk in real-time:
invoke(), which returns a single AIMessage after the model has finished generating its full response, stream() returns multiple AIMessageChunk objects, each containing a portion of the output text. Importantly, each chunk in a stream is designed to be gathered into a full message via summation:
invoke()—for example, it can be aggregated into a message history and passed back to the model as conversational context.
Advanced streaming topics
Advanced streaming topics
Streaming events
Streaming events
astream_events().This simplifies filtering based on event types and other metadata, and will aggregate the full message in the background. See below for an example."Auto-streaming" chat models
"Auto-streaming" chat models
model.invoke() within nodes, but LangChain will automatically delegate to streaming if running in a streaming mode.How it works
When youinvoke() a chat model, LangChain will automatically switch to an internal streaming mode if it detects that you are trying to stream the overall application. The result of the invocation will be the same as far as the code that was using invoke is concerned; however, while the chat model is being streamed, LangChain will take care of invoking on_llm_new_token events in LangChain’s callback system.Callback events allow LangGraph stream() and astream_events() to surface the chat model’s output in real-time.Batch
Batching a collection of independent requests to a model can significantly improve performance and reduce costs, as the processing can be done in parallel:batch() will only return the final output for the entire batch. If you want to receive the output for each individual input as it finishes generating, you can stream results with batch_as_completed():
batch_as_completed(), results may arrive out of order. Each includes the input index for matching to reconstruct the original order as needed.Tool calling
Models can request to call tools that perform tasks such as fetching data from a database, searching the web, or running code. Tools are pairings of:- A schema, including the name of the tool, a description, and/or argument definitions (often a JSON schema)
- A function or to execute.
bind_tools. In subsequent invocations, the model can choose to call any of the bound tools as needed.
Some model providers offer that can be enabled via model or invocation parameters (e.g. ChatOpenAI, ChatAnthropic). Check the respective provider reference for details.
Tool execution loop
Tool execution loop
ToolMessage returned by the tool includes a tool_call_id that matches the original tool call, helping the model correlate results with requests.Forcing tool calls
Forcing tool calls
Parallel tool calls
Parallel tool calls
Streaming tool calls
Streaming tool calls
ToolCallChunk. This allows you to see tool calls as they’re being generated rather than waiting for the complete response.Structured output
Models can be requested to provide their response in a format matching a given schema. This is useful for ensuring the output can be easily parsed and used in subsequent processing. LangChain supports multiple schema types and methods for enforcing structured output.- Pydantic
- TypedDict
- JSON Schema
- Method parameter: Some providers support different methods for structured output:
'json_schema': Uses dedicated structured output features offered by the provider.'function_calling': Derives structured output by forcing a tool call that follows the given schema.'json_mode': A precursor to'json_schema'offered by some providers. Generates valid JSON, but the schema must be described in the prompt.
- Include raw: Set
include_raw=Trueto get both the parsed output and the raw AI message. - Validation: Pydantic models provide automatic validation.
TypedDictand JSON Schema require manual validation.
Example: Message output alongside parsed structure
Example: Message output alongside parsed structure
AIMessage object alongside the parsed representation to access response metadata such as token counts. To do this, set include_raw=True when calling with_structured_output:Example: Nested structures
Example: Nested structures
Advanced topics
Model profiles
langchain>=1.1.profile attribute:
- Summarization middleware can trigger summarization based on a model’s context window size.
- Structured output strategies in
create_agentcan be inferred automatically (e.g., by checking support for native structured output features). - Model inputs can be gated based on supported modalities and maximum input tokens.
- Deep Agents Code filters the interactive model switcher to models whose profiles report
tool_callingsupport and text I/O, and displays context window sizes and capability flags in the selector detail view.
Updating or overwriting profile data
Updating or overwriting profile data
profile is also a regular dict and can be updated in place. If the model instance is shared, consider using model_copy to avoid mutating shared state.- (If needed) update the source data at models.dev through a pull request to its repository on GitHub.
- (If needed) update additional fields and overrides in
langchain_<package>/data/profile_augmentations.tomlthrough a pull request to the LangChain integration package`. - Use the
langchain-model-profilesCLI tool to pull the latest data from models.dev, merge in the augmentations and update the profile data:
- Downloads the latest data for
<provider>from models.dev - Merges augmentations from
profile_augmentations.tomlin<data_dir> - Writes merged profiles to
profiles.pyin<data_dir>
libs/partners/anthropic in the LangChain monorepo:Multimodal
Certain models can process and return non-textual data such as images, audio, and video. You can pass non-textual data to a model by providing content blocks. See the multimodal section of the messages guide for details. can return multimodal data as part of their response. If invoked to do so, the resultingAIMessage will have content blocks with multimodal types.
Reasoning
Many models are capable of performing multi-step reasoning to arrive at a conclusion. This involves breaking down complex problems into smaller, more manageable steps. If supported by the underlying model, you can surface this reasoning process to better understand how the model arrived at its final answer.'low' or 'high') or integer token budgets.
reasoning_effort as a standard parameter requires langchain-core>=1.5.2, plus the corresponding partner package version: langchain-anthropic>=1.5.3, langchain-openai>=1.4.1, langchain-fireworks>=1.5.2, langchain-xai>=1.3.0, langchain-google-genai>=4.3.1, or langchain-aws>=1.7.0.ChatOpenAI, ChatAnthropic, ChatFireworks, ChatXAI, ChatGoogleGenerativeAI, and ChatBedrockConverse support a standard reasoning_effort parameter. Like temperature, it can be set at model construction or per invocation, and each provider translates it into its own API format:
reasoning_effort (for example, ChatAnthropic accepts effort and ChatGoogleGenerativeAI accepts thinking_level). See the chat model integrations page for provider-specific detail.
For details, see the integrations page or reference for your respective chat model.
Local models
LangChain supports running models locally on your own hardware. This is useful for scenarios where either data privacy is critical, you want to invoke a custom model, or when you want to avoid the costs incurred when using a cloud-based model. Ollama is one of the easiest ways to run chat and embedding models locally.Prompt caching
Many providers offer prompt caching features to reduce latency and cost on repeat processing of the same tokens. You can engage caching at three levels:- Implicit provider caching: providers automatically pass on cost savings if a request hits a cache, with no configuration required. Examples: OpenAI and Gemini.
- Provider-level explicit controls: providers let you manually indicate cache points for greater control or to guarantee cost savings. These mirror the underlying provider/API behavior. Examples:
ChatOpenAI(viaprompt_cache_key)- Anthropic content-block
cache_control - Gemini.
- AWS Bedrock
cachePointblocks
- LangChain middleware: for agents, middleware lets LangChain optimize caching of stable system prompt and tool content. Examples:
- Anthropic’s
AnthropicPromptCachingMiddleware - AWS Bedrock’s
BedrockPromptCachingMiddleware
- Anthropic’s
Server-side tool use
Some providers support server-side tool-calling loops: models can interact with web search, code interpreters, and other tools and analyze the results in a single conversational turn. If a model invokes a tool server-side, the content of the response message will include content representing the invocation and result of the tool. Accessing the content blocks of the response will return the server-side tool calls and results in a provider-agnostic format:Model exceptions
Major integration packages raise standard exception types fromlangchain_core.exceptions for common model failures such as authentication errors, rate limits, and timeouts. These exceptions inherit from both the LangChain base type and the provider SDK’s own exception type, so you can catch either one:
is_retryable attribute that retry middleware respects by default.
Exception types
Exception types
ModelAuthenticationError— missing, invalid, or expired API key (not retryable)ModelPermissionDeniedError— credentials lack permission (not retryable)ModelInvalidRequestError— provider rejects the request (not retryable)ModelNotFoundError— requested model not found (not retryable)ModelRateLimitError— provider rate limit exceeded (retryable)ModelAPIError— provider server failure (retryable)ModelConnectionError— provider cannot be reached (retryable)ModelTimeoutError— request times out (retryable)ContextOverflowError— input exceeds the model context limit (not retryable)
Rate limiting
Many chat model providers impose a limit on the number of invocations that can be made in a given time period. If you hit a rate limit, you will typically receive a rate limit error response from the provider, and will need to wait before making more requests. To help manage rate limits, chat model integrations accept arate_limiter parameter that can be provided during initialization to control the rate at which requests are made.
Initialize and use a rate limiter
Initialize and use a rate limiter
InMemoryRateLimiter. This limiter is thread safe and can be shared by multiple threads in the same process.Base URL and proxy settings
You can configure a custom base URL for providers that implement the OpenAI Chat Completions API.Custom base URL
Custom base URL
init_chat_model with these providers by specifying the appropriate base_url parameter:HTTP proxy configuration
HTTP proxy configuration
Log probabilities
Certain models can be configured to return token-level log probabilities representing the likelihood of a given token by setting thelogprobs parameter when initializing the model:
Token usage
A number of model providers return token usage information as part of the invocation response. When available, this information will be included on theAIMessage objects produced by the corresponding model. For more details, see the messages guide.
- Callback handler
- Context manager
Invocation config
When invoking a model, you can pass additional configuration through theconfig parameter using a RunnableConfig dictionary. This provides run-time control over execution behavior, callbacks, and metadata tracking.
Common configuration options include:
- Debugging with LangSmith tracing
- Implementing custom logging or monitoring
- Controlling resource usage in production
- Tracking invocations across complex pipelines
Key configuration attributes
Key configuration attributes
batch() or batch_as_completed().Configurable models
You can also create a runtime-configurable model by specifyingconfigurable_fields. If you don’t specify a model value, then 'model' and 'model_provider' will be configurable by default.
Configurable model with default values
Configurable model with default values
init_chat_model reference for more details on configurable_fields and config_prefix.Using a configurable model declaratively
Using a configurable model declaratively
bind_tools, with_structured_output, with_configurable, etc. on a configurable model and chain a configurable model in the same way that we would a regularly instantiated chat model object.Dynamic model selection
Dynamic models are selected at based on the current and context. This enables sophisticated routing logic and cost optimization. To use a dynamic model, create middleware using the@wrap_model_call decorator that modifies the model in the request:

