codekingpro/portable-devtools
114k
1"""Standard, multimodal content blocks for Large Language Model I/O.2 3This module provides standardized data structures for representing inputs to and outputs4from LLMs. The core abstraction is the **Content Block**, a `TypedDict`.5 6**Rationale**7 8Different LLM providers use distinct and incompatible API schemas. This module provides9a unified, provider-agnostic format to facilitate these interactions. A message to or10from a model is simply a list of content blocks, allowing for the natural interleaving11of text, images, and other content in a single ordered sequence.12 13An adapter for a specific provider is responsible for translating this standard list of14blocks into the format required by its API.15 16**Extensibility**17 18Data **not yet mapped** to a standard block may be represented using the19`NonStandardContentBlock`, which allows for provider-specific data to be included20without losing the benefits of type checking and validation.21 22Furthermore, provider-specific fields **within** a standard block are fully supported23by default in the `extras` field of each block. This allows for additional metadata24to be included without breaking the standard structure. For example, Google's thought25signature:26 27```python28AIMessage(29 content=[30 {31 "type": "text",32 "text": "J'adore la programmation.",33 "extras": {"signature": "EpoWCpc..."}, # Thought signature34 }35 ], ...36)37```38 39 40!!! note41 42 Following widespread adoption of [PEP 728](https://peps.python.org/pep-0728/), we43 intend to add `extra_items=Any` as a param to Content Blocks. This will signify to44 type checkers that additional provider-specific fields are allowed outside of the45 `extras` field, and that will become the new standard approach to adding46 provider-specific metadata.47 48 ??? note49 50 **Example with PEP 728 provider-specific fields:**51 52 ```python53 # Content block definition54 # NOTE: `extra_items=Any`55 class TextContentBlock(TypedDict, extra_items=Any):56 type: Literal["text"]57 id: NotRequired[str]58 text: str59 annotations: NotRequired[list[Annotation]]60 index: NotRequired[int]61 ```62 63 ```python64 from langchain_core.messages.content import TextContentBlock65 66 # Create a text content block with provider-specific fields67 my_block: TextContentBlock = {68 # Add required fields69 "type": "text",70 "text": "Hello, world!",71 # Additional fields not specified in the TypedDict72 # These are valid with PEP 728 and are typed as Any73 "openai_metadata": {"model": "gpt-4", "temperature": 0.7},74 "anthropic_usage": {"input_tokens": 10, "output_tokens": 20},75 "custom_field": "any value",76 }77 78 # Mutating an existing block to add provider-specific fields79 openai_data = my_block["openai_metadata"] # Type: Any80 ```81 82**Example Usage**83 84```python85# Direct construction86from langchain_core.messages.content import TextContentBlock, ImageContentBlock87 88multimodal_message: AIMessage(89 content_blocks=[90 TextContentBlock(type="text", text="What is shown in this image?"),91 ImageContentBlock(92 type="image",93 url="https://www.langchain.com/images/brand/langchain_logo_text_w_white.png",94 mime_type="image/png",95 ),96 ]97)98 99# Using factories100from langchain_core.messages.content import create_text_block, create_image_block101 102multimodal_message: AIMessage(103 content=[104 create_text_block("What is shown in this image?"),105 create_image_block(106 url="https://www.langchain.com/images/brand/langchain_logo_text_w_white.png",107 mime_type="image/png",108 ),109 ]110)111```112 113Factory functions offer benefits such as:114 115- Automatic ID generation (when not provided)116- No need to manually specify the `type` field117"""118 119from typing import Any, Literal, get_args, get_type_hints120 121from typing_extensions import NotRequired, TypedDict122 123from langchain_core.utils.utils import ensure_id124 125 126class Citation(TypedDict):127 """Annotation for citing data from a document.128 129 !!! note130 131 `start`/`end` indices refer to the **response text**,132 not the source text. This means that the indices are relative to the model's133 response, not the original document (as specified in the `url`).134 135 !!! note "Factory function"136 137 `create_citation` may also be used as a factory to create a `Citation`.138 Benefits include:139 140 * Automatic ID generation (when not provided)141 * Required arguments strictly validated at creation time142 """143 144 type: Literal["citation"]145 """Type of the content block. Used for discrimination."""146 147 id: NotRequired[str]148 """Unique identifier for this content block.149 150 Either:151 152 - Generated by the provider153 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))154 """155 156 url: NotRequired[str]157 """URL of the document source."""158 159 title: NotRequired[str]160 """Source document title.161 162 For example, the page title for a web page or the title of a paper.163 """164 165 start_index: NotRequired[int]166 """Start index of the **response text** (`TextContentBlock.text`)."""167 168 end_index: NotRequired[int]169 """End index of the **response text** (`TextContentBlock.text`)"""170 171 cited_text: NotRequired[str]172 """Excerpt of source text being cited."""173 174 # NOTE: not including spans for the raw document text (such as `text_start_index`175 # and `text_end_index`) as this is not currently supported by any provider. The176 # thinking is that the `cited_text` should be sufficient for most use cases, and it177 # is difficult to reliably extract spans from the raw document text across file178 # formats or encoding schemes.179 180 extras: NotRequired[dict[str, Any]]181 """Provider-specific metadata."""182 183 184class NonStandardAnnotation(TypedDict):185 """Provider-specific annotation format."""186 187 type: Literal["non_standard_annotation"]188 """Type of the content block. Used for discrimination."""189 190 id: NotRequired[str]191 """Unique identifier for this content block.192 193 Either:194 195 - Generated by the provider196 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))197 """198 199 value: dict[str, Any]200 """Provider-specific annotation data."""201 202 203Annotation = Citation | NonStandardAnnotation204"""A union of all defined `Annotation` types."""205 206 207class TextContentBlock(TypedDict):208 """Text output from a LLM.209 210 This typically represents the main text content of a message, such as the response211 from a language model or the text of a user message.212 213 !!! note "Factory function"214 215 `create_text_block` may also be used as a factory to create a216 `TextContentBlock`. Benefits include:217 218 * Automatic ID generation (when not provided)219 * Required arguments strictly validated at creation time220 """221 222 type: Literal["text"]223 """Type of the content block. Used for discrimination."""224 225 id: NotRequired[str]226 """Unique identifier for this content block.227 228 Either:229 230 - Generated by the provider231 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))232 """233 234 text: str235 """Block text."""236 237 annotations: NotRequired[list[Annotation]]238 """`Citation`s and other annotations."""239 240 index: NotRequired[int | str]241 """Index of block in aggregate response. Used during streaming."""242 243 extras: NotRequired[dict[str, Any]]244 """Provider-specific metadata."""245 246 247class ToolCall(TypedDict):248 """Represents an AI's request to call a tool.249 250 Example:251 ```python252 {"name": "foo", "args": {"a": 1}, "id": "123"}253 ```254 255 This represents a request to call the tool named "foo" with arguments {"a": 1}256 and an identifier of "123".257 258 !!! note "Factory function"259 260 `create_tool_call` may also be used as a factory to create a261 `ToolCall`. Benefits include:262 263 * Automatic ID generation (when not provided)264 * Required arguments strictly validated at creation time265 """266 267 type: Literal["tool_call"]268 """Used for discrimination."""269 270 id: str | None271 """An identifier associated with the tool call.272 273 An identifier is needed to associate a tool call request with a tool274 call result in events when multiple concurrent tool calls are made.275 """276 # TODO: Consider making this NotRequired[str] in the future.277 278 name: str279 """The name of the tool to be called."""280 281 args: dict[str, Any]282 """The arguments to the tool call."""283 284 index: NotRequired[int | str]285 """Index of block in aggregate response. Used during streaming."""286 287 extras: NotRequired[dict[str, Any]]288 """Provider-specific metadata."""289 290 291class ToolCallChunk(TypedDict):292 """A chunk of a tool call (yielded when streaming).293 294 When merging `ToolCallChunks` (e.g., via `AIMessageChunk.__add__`),295 all string attributes are concatenated. Chunks are only merged if their296 values of `index` are equal and not `None`.297 298 Example:299 ```python300 left_chunks = [ToolCallChunk(name="foo", args='{"a":', index=0)]301 right_chunks = [ToolCallChunk(name=None, args="1}", index=0)]302 303 (304 AIMessageChunk(content="", tool_call_chunks=left_chunks)305 + AIMessageChunk(content="", tool_call_chunks=right_chunks)306 ).tool_call_chunks == [ToolCallChunk(name="foo", args='{"a":1}', index=0)]307 ```308 """309 310 # TODO: Consider making fields NotRequired[str] in the future.311 312 type: Literal["tool_call_chunk"]313 """Used for serialization."""314 315 id: str | None316 """An identifier associated with the tool call.317 318 An identifier is needed to associate a tool call request with a tool319 call result in events when multiple concurrent tool calls are made.320 """321 # TODO: Consider making this NotRequired[str] in the future.322 323 name: str | None324 """The name of the tool to be called."""325 326 args: str | None327 """The arguments to the tool call."""328 329 index: NotRequired[int | str]330 """The index of the tool call in a sequence."""331 332 extras: NotRequired[dict[str, Any]]333 """Provider-specific metadata."""334 335 336class InvalidToolCall(TypedDict):337 """Allowance for errors made by LLM.338 339 Here we add an `error` key to surface errors made during generation340 (e.g., invalid JSON arguments.)341 """342 343 # TODO: Consider making fields NotRequired[str] in the future.344 345 type: Literal["invalid_tool_call"]346 """Used for discrimination."""347 348 id: str | None349 """An identifier associated with the tool call.350 351 An identifier is needed to associate a tool call request with a tool352 call result in events when multiple concurrent tool calls are made.353 """354 # TODO: Consider making this NotRequired[str] in the future.355 356 name: str | None357 """The name of the tool to be called."""358 359 args: str | None360 """The arguments to the tool call."""361 362 error: str | None363 """An error message associated with the tool call."""364 365 index: NotRequired[int | str]366 """Index of block in aggregate response. Used during streaming."""367 368 extras: NotRequired[dict[str, Any]]369 """Provider-specific metadata."""370 371 372class ServerToolCall(TypedDict):373 """Tool call that is executed server-side.374 375 For example: code execution, web search, etc.376 """377 378 type: Literal["server_tool_call"]379 """Used for discrimination."""380 381 id: str382 """An identifier associated with the tool call."""383 384 name: str385 """The name of the tool to be called."""386 387 args: dict[str, Any]388 """The arguments to the tool call."""389 390 index: NotRequired[int | str]391 """Index of block in aggregate response. Used during streaming."""392 393 extras: NotRequired[dict[str, Any]]394 """Provider-specific metadata."""395 396 397class ServerToolCallChunk(TypedDict):398 """A chunk of a server-side tool call (yielded when streaming)."""399 400 type: Literal["server_tool_call_chunk"]401 """Used for discrimination."""402 403 name: NotRequired[str]404 """The name of the tool to be called."""405 406 args: NotRequired[str]407 """JSON substring of the arguments to the tool call."""408 409 id: NotRequired[str]410 """Unique identifier for this server tool call chunk.411 412 Either:413 414 - Generated by the provider415 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))416 """417 418 index: NotRequired[int | str]419 """Index of block in aggregate response. Used during streaming."""420 421 extras: NotRequired[dict[str, Any]]422 """Provider-specific metadata."""423 424 425class ServerToolResult(TypedDict):426 """Result of a server-side tool call."""427 428 type: Literal["server_tool_result"]429 """Used for discrimination."""430 431 id: NotRequired[str]432 """Unique identifier for this server tool result.433 434 Either:435 436 - Generated by the provider437 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))438 """439 440 tool_call_id: str441 """ID of the corresponding server tool call."""442 443 status: Literal["success", "error"]444 """Execution status of the server-side tool."""445 446 output: NotRequired[Any]447 """Output of the executed tool."""448 449 index: NotRequired[int | str]450 """Index of block in aggregate response. Used during streaming."""451 452 extras: NotRequired[dict[str, Any]]453 """Provider-specific metadata."""454 455 456class ReasoningContentBlock(TypedDict):457 """Reasoning output from a LLM.458 459 !!! note "Factory function"460 461 `create_reasoning_block` may also be used as a factory to create a462 `ReasoningContentBlock`. Benefits include:463 464 * Automatic ID generation (when not provided)465 * Required arguments strictly validated at creation time466 """467 468 type: Literal["reasoning"]469 """Type of the content block. Used for discrimination."""470 471 id: NotRequired[str]472 """Unique identifier for this content block.473 474 Either:475 476 - Generated by the provider477 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))478 """479 480 reasoning: NotRequired[str]481 """Reasoning text.482 483 Either the thought summary or the raw reasoning text itself.484 485 Often parsed from `<think>` tags in the model's response.486 """487 488 index: NotRequired[int | str]489 """Index of block in aggregate response. Used during streaming."""490 491 extras: NotRequired[dict[str, Any]]492 """Provider-specific metadata."""493 494 495# Note: `title` and `context` are fields that could be used to provide additional496# information about the file, such as a description or summary of its content.497# E.g. with Claude, you can provide a context for a file which is passed to the model.498class ImageContentBlock(TypedDict):499 """Image data.500 501 !!! note "Factory function"502 503 `create_image_block` may also be used as a factory to create an504 `ImageContentBlock`. Benefits include:505 506 * Automatic ID generation (when not provided)507 * Required arguments strictly validated at creation time508 """509 510 type: Literal["image"]511 """Type of the content block. Used for discrimination."""512 513 id: NotRequired[str]514 """Unique identifier for this content block.515 516 Either:517 518 - Generated by the provider519 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))520 """521 522 file_id: NotRequired[str]523 """Reference to the image in an external file storage system.524 525 For example, OpenAI or Anthropic's Files API.526 """527 528 mime_type: NotRequired[str]529 """MIME type of the image.530 531 Required for base64 data.532 533 [Examples from IANA](https://www.iana.org/assignments/media-types/media-types.xhtml#image)534 """535 536 index: NotRequired[int | str]537 """Index of block in aggregate response. Used during streaming."""538 539 url: NotRequired[str]540 """URL of the image."""541 542 base64: NotRequired[str]543 """Data as a base64 string."""544 545 extras: NotRequired[dict[str, Any]]546 """Provider-specific metadata. This shouldn't be used for the image data itself."""547 548 549class VideoContentBlock(TypedDict):550 """Video data.551 552 !!! note "Factory function"553 554 `create_video_block` may also be used as a factory to create a555 `VideoContentBlock`. Benefits include:556 557 * Automatic ID generation (when not provided)558 * Required arguments strictly validated at creation time559 """560 561 type: Literal["video"]562 """Type of the content block. Used for discrimination."""563 564 id: NotRequired[str]565 """Unique identifier for this content block.566 567 Either:568 569 - Generated by the provider570 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))571 """572 573 file_id: NotRequired[str]574 """Reference to the video in an external file storage system.575 576 For example, OpenAI or Anthropic's Files API.577 """578 579 mime_type: NotRequired[str]580 """MIME type of the video.581 582 Required for base64 data.583 584 [Examples from IANA](https://www.iana.org/assignments/media-types/media-types.xhtml#video)585 """586 587 index: NotRequired[int | str]588 """Index of block in aggregate response. Used during streaming."""589 590 url: NotRequired[str]591 """URL of the video."""592 593 base64: NotRequired[str]594 """Data as a base64 string."""595 596 extras: NotRequired[dict[str, Any]]597 """Provider-specific metadata. This shouldn't be used for the video data itself."""598 599 600class AudioContentBlock(TypedDict):601 """Audio data.602 603 !!! note "Factory function"604 605 `create_audio_block` may also be used as a factory to create an606 `AudioContentBlock`. Benefits include:607 608 * Automatic ID generation (when not provided)609 * Required arguments strictly validated at creation time610 """611 612 type: Literal["audio"]613 """Type of the content block. Used for discrimination."""614 615 id: NotRequired[str]616 """Unique identifier for this content block.617 618 Either:619 620 - Generated by the provider621 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))622 """623 624 file_id: NotRequired[str]625 """Reference to the audio file in an external file storage system.626 627 For example, OpenAI or Anthropic's Files API.628 """629 630 mime_type: NotRequired[str]631 """MIME type of the audio.632 633 Required for base64 data.634 635 [Examples from IANA](https://www.iana.org/assignments/media-types/media-types.xhtml#audio)636 """637 638 index: NotRequired[int | str]639 """Index of block in aggregate response. Used during streaming."""640 641 url: NotRequired[str]642 """URL of the audio."""643 644 base64: NotRequired[str]645 """Data as a base64 string."""646 647 extras: NotRequired[dict[str, Any]]648 """Provider-specific metadata. This shouldn't be used for the audio data itself."""649 650 651class PlainTextContentBlock(TypedDict):652 """Plaintext data (e.g., from a `.txt` or `.md` document).653 654 !!! note655 656 A `PlainTextContentBlock` existed in `langchain-core<1.0.0`. Although the657 name has carried over, the structure has changed significantly. The only shared658 keys between the old and new versions are `type` and `text`, though the659 `type` value has changed from `'text'` to `'text-plain'`.660 661 !!! note662 663 Title and context are optional fields that may be passed to the model. See664 Anthropic [example](https://platform.claude.com/docs/en/build-with-claude/citations#citable-vs-non-citable-content).665 666 !!! note "Factory function"667 668 `create_plaintext_block` may also be used as a factory to create a669 `PlainTextContentBlock`. Benefits include:670 671 * Automatic ID generation (when not provided)672 * Required arguments strictly validated at creation time673 """674 675 type: Literal["text-plain"]676 """Type of the content block. Used for discrimination."""677 678 id: NotRequired[str]679 """Unique identifier for this content block.680 681 Either:682 683 - Generated by the provider684 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))685 """686 687 file_id: NotRequired[str]688 """Reference to the plaintext file in an external file storage system.689 690 For example, OpenAI or Anthropic's Files API.691 """692 693 mime_type: Literal["text/plain"]694 """MIME type of the file.695 696 Required for base64 data.697 """698 699 index: NotRequired[int | str]700 """Index of block in aggregate response. Used during streaming."""701 702 url: NotRequired[str]703 """URL of the plaintext."""704 705 base64: NotRequired[str]706 """Data as a base64 string."""707 708 text: NotRequired[str]709 """Plaintext content. This is optional if the data is provided as base64."""710 711 title: NotRequired[str]712 """Title of the text data, e.g., the title of a document."""713 714 context: NotRequired[str]715 """Context for the text, e.g., a description or summary of the text's content."""716 717 extras: NotRequired[dict[str, Any]]718 """Provider-specific metadata. This shouldn't be used for the data itself."""719 720 721class FileContentBlock(TypedDict):722 """File data that doesn't fit into other multimodal block types.723 724 This block is intended for files that are not images, audio, or plaintext. For725 example, it can be used for PDFs, Word documents, etc.726 727 If the file is an image, audio, or plaintext, you should use the corresponding728 content block type (e.g., `ImageContentBlock`, `AudioContentBlock`,729 `PlainTextContentBlock`).730 731 !!! note "Factory function"732 733 `create_file_block` may also be used as a factory to create a734 `FileContentBlock`. Benefits include:735 736 * Automatic ID generation (when not provided)737 * Required arguments strictly validated at creation time738 """739 740 type: Literal["file"]741 """Type of the content block. Used for discrimination."""742 743 id: NotRequired[str]744 """Unique identifier for this content block.745 746 Used for tracking and referencing specific blocks (e.g., during streaming).747 748 Not to be confused with `file_id`, which references an external file in a749 storage system.750 751 Either:752 753 - Generated by the provider754 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))755 """756 757 file_id: NotRequired[str]758 """Reference to the file in an external file storage system.759 760 For example, a file ID from OpenAI's Files API or another cloud storage provider.761 This is distinct from `id`, which identifies the content block itself.762 """763 764 mime_type: NotRequired[str]765 """MIME type of the file.766 767 Required for base64 data.768 769 [Examples from IANA](https://www.iana.org/assignments/media-types/media-types.xhtml)770 """771 772 index: NotRequired[int | str]773 """Index of block in aggregate response. Used during streaming."""774 775 url: NotRequired[str]776 """URL of the file."""777 778 base64: NotRequired[str]779 """Data as a base64 string."""780 781 extras: NotRequired[dict[str, Any]]782 """Provider-specific metadata. This shouldn't be used for the file data itself."""783 784 785# Future modalities to consider:786# - 3D models787# - Tabular data788 789 790class NonStandardContentBlock(TypedDict):791 """Provider-specific content data.792 793 This block contains data for which there is not yet a standard type.794 795 The purpose of this block should be to simply hold a provider-specific payload.796 If a provider's non-standard output includes reasoning and tool calls, it should be797 the adapter's job to parse that payload and emit the corresponding standard798 `ReasoningContentBlock` and `ToolCalls`.799 800 Has no `extras` field, as provider-specific data should be included in the801 `value` field.802 803 !!! note "Factory function"804 805 `create_non_standard_block` may also be used as a factory to create a806 `NonStandardContentBlock`. Benefits include:807 808 * Automatic ID generation (when not provided)809 * Required arguments strictly validated at creation time810 """811 812 type: Literal["non_standard"]813 """Type of the content block. Used for discrimination."""814 815 id: NotRequired[str]816 """Unique identifier for this content block.817 818 Either:819 820 - Generated by the provider821 - Generated by LangChain upon creation (`UUID4` prefixed with `'lc_'`))822 """823 824 value: dict[str, Any]825 """Provider-specific content data."""826 827 index: NotRequired[int | str]828 """Index of block in aggregate response. Used during streaming."""829 830 831# --- Aliases ---832DataContentBlock = (833 ImageContentBlock834 | VideoContentBlock835 | AudioContentBlock836 | PlainTextContentBlock837 | FileContentBlock838)839"""A union of all defined multimodal data `ContentBlock` types."""840 841ToolContentBlock = (842 ToolCall | ToolCallChunk | ServerToolCall | ServerToolCallChunk | ServerToolResult843)844 845ContentBlock = (846 TextContentBlock847 | InvalidToolCall848 | ReasoningContentBlock849 | NonStandardContentBlock850 | DataContentBlock851 | ToolContentBlock852)853"""A union of all defined `ContentBlock` types and aliases."""854 855 856KNOWN_BLOCK_TYPES = {857 # Text output858 "text",859 "reasoning",860 # Tools861 "tool_call",862 "invalid_tool_call",863 "tool_call_chunk",864 # Multimodal data865 "image",866 "audio",867 "file",868 "text-plain",869 "video",870 # Server-side tool calls871 "server_tool_call",872 "server_tool_call_chunk",873 "server_tool_result",874 # Catch-all875 "non_standard",876 # citation and non_standard_annotation intentionally omitted877}878"""These are block types known to `langchain-core >= 1.0.0`.879 880If a block has a type not in this set, it is considered to be provider-specific.881"""882 883 884def _get_data_content_block_types() -> tuple[str, ...]:885 """Get type literals from DataContentBlock union members dynamically.886 887 Example: ("image", "video", "audio", "text-plain", "file")888 889 Note that old style multimodal blocks type literals with new style blocks.890 Specifically, "image", "audio", and "file".891 892 See the docstring of `_normalize_messages` in `language_models._utils` for details.893 """894 data_block_types = []895 896 for block_type in get_args(DataContentBlock):897 hints = get_type_hints(block_type)898 if "type" in hints:899 type_annotation = hints["type"]900 if hasattr(type_annotation, "__args__"):901 # This is a Literal type, get the literal value902 literal_value = type_annotation.__args__[0]903 data_block_types.append(literal_value)904 905 return tuple(data_block_types)906 907 908def is_data_content_block(block: dict) -> bool:909 """Check if the provided content block is a data content block.910 911 Returns True for both v0 (old-style) and v1 (new-style) multimodal data blocks.912 913 Args:914 block: The content block to check.915 916 Returns:917 `True` if the content block is a data content block, `False` otherwise.918 """919 if block.get("type") not in _get_data_content_block_types():920 return False921 922 if any(key in block for key in ("url", "base64", "file_id", "text")):923 # Type is valid and at least one data field is present924 # (Accepts old-style image and audio URLContentBlock)925 926 # 'text' is checked to support v0 PlainTextContentBlock types927 # We must guard against new style TextContentBlock which also has 'text' `type`928 # by ensuring the presence of `source_type`929 if block["type"] == "text" and "source_type" not in block: # noqa: SIM103 # This is more readable930 return False931 932 return True933 934 if "source_type" in block:935 # Old-style content blocks had possible types of 'image', 'audio', and 'file'936 # which is not captured in the prior check937 source_type = block["source_type"]938 if (source_type == "url" and "url" in block) or (939 source_type == "base64" and "data" in block940 ):941 return True942 if (source_type == "id" and "id" in block) or (943 source_type == "text" and "url" in block944 ):945 return True946 947 return False948 949 950def create_text_block(951 text: str,952 *,953 id: str | None = None,954 annotations: list[Annotation] | None = None,955 index: int | str | None = None,956 **kwargs: Any,957) -> TextContentBlock:958 """Create a `TextContentBlock`.959 960 Args:961 text: The text content of the block.962 id: Content block identifier.963 964 Generated automatically if not provided.965 annotations: `Citation`s and other annotations for the text.966 index: Index of block in aggregate response.967 968 Used during streaming.969 970 Returns:971 A properly formatted `TextContentBlock`.972 973 !!! note974 975 The `id` is generated automatically if not provided, using a UUID4 format976 prefixed with `'lc_'` to indicate it is a LangChain-generated ID.977 """978 block = TextContentBlock(979 type="text",980 text=text,981 id=ensure_id(id),982 )983 if annotations is not None:984 block["annotations"] = annotations985 if index is not None:986 block["index"] = index987 988 extras = {k: v for k, v in kwargs.items() if v is not None}989 if extras:990 block["extras"] = extras991 992 return block993 994 995def create_image_block(996 *,997 url: str | None = None,998 base64: str | None = None,999 file_id: str | None = None,1000 mime_type: str | None = None,1001 id: str | None = None,1002 index: int | str | None = None,1003 **kwargs: Any,1004) -> ImageContentBlock:1005 """Create an `ImageContentBlock`.1006 1007 Args:1008 url: URL of the image.1009 base64: Base64-encoded image data.1010 file_id: ID of the image file from a file storage system.1011 mime_type: MIME type of the image.1012 1013 Required for base64 data.1014 id: Content block identifier.1015 1016 Generated automatically if not provided.1017 index: Index of block in aggregate response.1018 1019 Used during streaming.1020 1021 Returns:1022 A properly formatted `ImageContentBlock`.1023 1024 Raises:1025 ValueError: If no image source is provided or if `base64` is used without1026 `mime_type`.1027 1028 !!! note1029 1030 The `id` is generated automatically if not provided, using a UUID4 format1031 prefixed with `'lc_'` to indicate it is a LangChain-generated ID.1032 """1033 if not any([url, base64, file_id]):1034 msg = "Must provide one of: url, base64, or file_id"1035 raise ValueError(msg)1036 1037 block = ImageContentBlock(type="image", id=ensure_id(id))1038 1039 if url is not None:1040 block["url"] = url1041 if base64 is not None:1042 block["base64"] = base641043 if file_id is not None:1044 block["file_id"] = file_id1045 if mime_type is not None:1046 block["mime_type"] = mime_type1047 if index is not None:1048 block["index"] = index1049 1050 extras = {k: v for k, v in kwargs.items() if v is not None}1051 if extras:1052 block["extras"] = extras1053 1054 return block1055 1056 1057def create_video_block(1058 *,1059 url: str | None = None,1060 base64: str | None = None,1061 file_id: str | None = None,1062 mime_type: str | None = None,1063 id: str | None = None,1064 index: int | str | None = None,1065 **kwargs: Any,1066) -> VideoContentBlock:1067 """Create a `VideoContentBlock`.1068 1069 Args:1070 url: URL of the video.1071 base64: Base64-encoded video data.1072 file_id: ID of the video file from a file storage system.1073 mime_type: MIME type of the video.1074 1075 Required for base64 data.1076 id: Content block identifier.1077 1078 Generated automatically if not provided.1079 index: Index of block in aggregate response.1080 1081 Used during streaming.1082 1083 Returns:1084 A properly formatted `VideoContentBlock`.1085 1086 Raises:1087 ValueError: If no video source is provided or if `base64` is used without1088 `mime_type`.1089 1090 !!! note1091 1092 The `id` is generated automatically if not provided, using a UUID4 format1093 prefixed with `'lc_'` to indicate it is a LangChain-generated ID.1094 """1095 if not any([url, base64, file_id]):1096 msg = "Must provide one of: url, base64, or file_id"1097 raise ValueError(msg)1098 1099 if base64 and not mime_type:1100 msg = "mime_type is required when using base64 data"1101 raise ValueError(msg)1102 1103 block = VideoContentBlock(type="video", id=ensure_id(id))1104 1105 if url is not None:1106 block["url"] = url1107 if base64 is not None:1108 block["base64"] = base641109 if file_id is not None:1110 block["file_id"] = file_id1111 if mime_type is not None:1112 block["mime_type"] = mime_type1113 if index is not None:1114 block["index"] = index1115 1116 extras = {k: v for k, v in kwargs.items() if v is not None}1117 if extras:1118 block["extras"] = extras1119 1120 return block1121 1122 1123def create_audio_block(1124 *,1125 url: str | None = None,1126 base64: str | None = None,1127 file_id: str | None = None,1128 mime_type: str | None = None,1129 id: str | None = None,1130 index: int | str | None = None,1131 **kwargs: Any,1132) -> AudioContentBlock:1133 """Create an `AudioContentBlock`.1134 1135 Args:1136 url: URL of the audio.1137 base64: Base64-encoded audio data.1138 file_id: ID of the audio file from a file storage system.1139 mime_type: MIME type of the audio.1140 1141 Required for base64 data.1142 id: Content block identifier.1143 1144 Generated automatically if not provided.1145 index: Index of block in aggregate response.1146 1147 Used during streaming.1148 1149 Returns:1150 A properly formatted `AudioContentBlock`.1151 1152 Raises:1153 ValueError: If no audio source is provided or if `base64` is used without1154 `mime_type`.1155 1156 !!! note1157 1158 The `id` is generated automatically if not provided, using a UUID4 format1159 prefixed with `'lc_'` to indicate it is a LangChain-generated ID.1160 """1161 if not any([url, base64, file_id]):1162 msg = "Must provide one of: url, base64, or file_id"1163 raise ValueError(msg)1164 1165 if base64 and not mime_type:1166 msg = "mime_type is required when using base64 data"1167 raise ValueError(msg)1168 1169 block = AudioContentBlock(type="audio", id=ensure_id(id))1170 1171 if url is not None:1172 block["url"] = url1173 if base64 is not None:1174 block["base64"] = base641175 if file_id is not None:1176 block["file_id"] = file_id1177 if mime_type is not None:1178 block["mime_type"] = mime_type1179 if index is not None:1180 block["index"] = index1181 1182 extras = {k: v for k, v in kwargs.items() if v is not None}1183 if extras:1184 block["extras"] = extras1185 1186 return block1187 1188 1189def create_file_block(1190 *,1191 url: str | None = None,1192 base64: str | None = None,1193 file_id: str | None = None,1194 mime_type: str | None = None,1195 id: str | None = None,1196 index: int | str | None = None,1197 **kwargs: Any,1198) -> FileContentBlock:1199 """Create a `FileContentBlock`.1200 