You need to enable JavaScript to run this app.
Lake AI Service

Lake AI Service

Copy page
Download PDF
Operators overview
Supported operators: online operators
Copy page
Download PDF
Supported operators: online operators

Video

Video generation

Operator name

Usage

Operator overview

Enhanced and basic versions of video generation
Seedance video generation (Doubao series)

  • Online
  • The LAS video generation enhanced/basic versions are built around the Seedance series models, focusing on the pre- and post-processing workflows of video generation. They integrate capabilities such as video format conversion, video segmentation, video restoration, subtitle generation, audio-video merging, and video super-resolution to extend the capabilities of video generation models, improve the quality and duration of generated videos, reduce token consumption, and improve production efficiency for customers.

Short drama narration

  • Online
  • Based on a single short drama or film material video, the short drama material commentary operator can automatically understand the plot content, recognize original video dialogue, generate commentary scripts in viral styles, complete commentary voiceover and video rendering, and ultimately output a new video with narrated commentary. The operator retains appropriate silent gaps in key plot segments and lowers the original audio volume when the commentary overlays the original audio, helping to quickly generate plot commentary material suitable for distribution on short video platforms.
  • Key features:
    • Video content understanding: Automatically analyzes the input video frames, plot progression, character relationships, and key events to provide a contextual basis for commentary script generation.
    • Original video ASR recognition: Automatically recognizes dialogue and audio content in the original video, assisting in determining plot information and highlighting original audio segments that can be retained.
    • Viral commentary script generation: Supports controlling commentary style via style_prompt, and by default generates short drama commentary scripts with a tight rhythm, hook-first approach, and avoids spoiling key plot twists.
    • Silence density control: Supports four levels of silence density—none, low, medium, and high—to control the proportion of commentary overlay and the degree of highlight retention from the original video.
    • TTS voiceover and time alignment: Automatically synthesizes the commentary script into speech and aligns it with the video time window, ensuring the narration rhythm matches the visual content; currently supports generating Chinese commentary voiceover, using the default preset voice.
    • Subtitle and video rendering: Automatically generates commentary subtitles and completes video rendering, outputting the final video with narration and subtitles.

Video Editing Enhanced Edition

  • Online
  • Enhanced online video editing service that replaces video content based on input video and reference images. Supports scene replacement and object replacement, strives to preserve the main actions, camera rhythm, and temporal continuity of the original video, and outputs the resulting video after replacement.
  • Key features:
    • Supports scene_replace for scene replacement, preserving people, products, actions, and camera structure in the original video while replacing the background or environmental style.
    • Supports object_replace for object replacement, replacing target objects in the video with specified styles or products based on reference images.
    • Supports supplementing natural language editing constraints via user_prompt, such as retaining actions, removing watermarks, avoiding subtitles, or controlling style details.
    • Supports automatic evaluation and retry for a limited number of times to help achieve more stable replacement results with complex materials.
    • Supports returning the final video result and the result.json diagnostic file, enabling callers to retain generated results and process information.

Video production

Operator name

Usage

Operator overview

Short drama script generation

  • Online
  • The short drama/movie script generation operator is an automated script reverse engineering tool designed for short dramas as well as long-form or serialized video content such as movies. The operator leverages a visual multimodal large model (VLM) to automatically extract all characters from the entire drama or film, analyze character relationships, and generate high-quality textual scripts and character lists with scene, action, expression, and dialogue details based on visuals and dialogue, supporting secondary creation, overseas translation, and copyright protection for video content.
  • Key features:
    • Consistent character identification: Overcomes the limitations of isolated understanding in single episodes or segments, enabling stable tracking of main characters in long-running series or feature-length films. Ensures a high degree of consistency in character identity and settings across episodes, even in cases of costume changes, side profiles, or complex scene transitions, ultimately building a complete global character list.
    • High-fidelity script reconstruction: Combines video visuals and dialogue to generate professional-grade storyboard scripts (with precise timestamps in movie mode). Accurately restores scene layout, character emotions, body movements, and key dialogue, providing high-quality textual drafts ready for secondary development or for translation or proofreading.
    • Dual-mode adaptive architecture:
      • Short drama mode: Supports batch input of multiple short drama episodes, processes them strictly in input order, and maintains continuity of the serialized plot and character consistency.
      • Movie mode: For movies or long recordings lasting several hours per episode, automatically initiates adaptive processing strategies for long videos, effectively alleviating the issue of detail loss caused by large models with long context windows.
    • Flexible output format customization: Provides an open custom instruction (Prompt) interface. You can freely adjust the script generation style for each episode according to specific business requirements (such as focusing on psychological descriptions, specific storyboard layout formats, specific text markers, and more), meeting the direct integration needs of various downstream businesses.
    • Convenient result delivery: Supports securely writing the generated character list and full script directly to your designated cloud storage (TOS), and can also generate packaged pre-signed download links.

Short drama script parsing

  • Online
  • The short drama script parsing operator automatically converts short drama script text into a structured asset table. The operator reads one or more script files, automatically identifies the script format and output language, extracts three types of assets—characters, props, and scenes—and adds visual details to character appearance, character styling, prop visual descriptions, and scene spatial descriptions. The final results are saved as JSON files.
  • Key features:
    • Multi-script input parsing: Supports input of one or more script files, concatenates and processes them in input order, suitable for parsing single episodes, multiple episodes, or entire series scripts.
    • Character asset table extraction: Identifies character names, gender, basic appearance, and the scenes in which the character appears, and provides multiple styling descriptions for main characters that match the script background.
    • Prop asset table extraction: Identifies key props, their locations, their role in the plot, and visual features to assist prop design and asset management.
    • Scene asset table extraction: Identifies main scenes, space types, time setting, atmosphere, and visual layout to assist scene construction and storyboard generation.
    • Chinese and English script adaptation: Supports Chinese short drama numbering format and English Hollywood/Fountain slug-line format; in auto mode, automatically identifies the script format and determines the asset table output language.
    • Stable processing of ultra-long scripts: Automatically segments scripts exceeding the length threshold along scene boundaries to reduce timeout and truncation risks caused by ultra-long context requests.
    • TOS deliverables: The character table, prop table, scene table, and script background information are all saved as JSON files, and the corresponding TOS paths are returned.

Short drama script fission

  • Online
  • Designed to perform targeted rewriting based on existing episodic scripts. The operator preserves the original script's core character structure, conflict chain, emotional arc, and key pacing points, while systematically rewriting character identities, world settings, regional context, industry background, and key terminology to generate a new script version tailored to the target market and subject matter. The output includes the directory of the newly split episodic scripts, the complete script, outline, and before-and-after mapping table, making it suitable for short drama overseas distribution, replicating hit shows, and testing content directions.
  • Key features:
    • Targeted rewriting with skeleton retention: Preserves the original script's character functions, character relationship structure, sequence of key events, and main storyline logic, while only rewriting surface settings and the context of expression.
    • Localization rewriting: Supports generating scripts for different regional contexts, including simultaneous adaptation of character naming styles, institutional backgrounds, dialogue expressions, and social relationship logic.
    • Subject-based rewriting: Supports overall style migration according to the target subject matter, rewriting the same plot skeleton into new scripts reusable across different content genres.
    • Batch processing of episodic scripts: Supports submitting multiple episodic scripts at once via directory or list, processes them in input order, and maintains narrative coherence across the complete script.
    • Structured delivery of results: Outputs the split episodic scripts, complete script, outline, and mapping table after transformation, facilitating subsequent video production, proofreading, manual refinement, and effectiveness review.

Video script planning

  • Online
  • Enter the video topic and creative requirements to generate a video production plan for subsequent production. The plan breaks down the creative idea into a clear content structure, shot arrangement, material suggestions, and timing and pacing, making it suitable for early-stage planning of educational and marketing videos.
  • Key features:
    • From creative idea to production plan: Organizes the topic, audience, and communication objectives into an actionable video production plan.
    • Shot and pacing planning: Break down video content and plan the shot direction, duration, and transitions for each section.
    • Material usage suggestions: Arrange the use of video, images, and other materials according to content requirements to facilitate subsequent production.
    • Support for different content types: Supports planning for course instruction and marketing or e-commerce videos.

Video scene segmentation

  • Online
  • The video scene segmentation operator uses a multimodal large model to segment input videos into shots/scenes, perform global character recognition, associate characters at the scene level, and extract character images. The operator outputs scene summary results, a character registry, video clips for each scene, and image files organized by character for subsequent retrieval, editing, and content understanding.
  • Key features:
    • Supports scene segmentation based on VLM, as well as equal-duration segmentation when min_segment_duration == max_segment_duration.
    • Supports global character extraction and deduplication aggregation to generate a character registry.
    • Supports associating characters within scenes, outputting the time intervals of character appearances, key frame timestamps, and bbox information in the scene.
    • Supports automatically extracting independent video files for each scene.
    • Supports extracting and filtering representative images for each character, and outputs them, organized by character.
    • Supports outputting token usage and LLM request counts to facilitate cost assessment.

Video understanding

Operator name

Usage

Operator overview

Fine-grained video understanding

  • Online
  • The LAS Video Fine-Grained Understanding API targets various types of video content, providing multidimensional and fine-grained structured understanding. Whether it is a short video, movie clip, or long meeting recording, users can upload videos to obtain searchable, Q&A-enabled content data and detailed summaries.
  • Key features
    • Global fine-grained understanding: Supports videos up to several hours (maximum 3h, 10G), generating coherent timelines and chapter summaries.
    • Event and action recognition: Accurately detects key events, character actions, scene changes, and logical relationships.
    • Video question answering: Natural language question answering based on video content, quickly locating answers and timestamps.
    • Efficient summarization and tagging: Automatically generates chapter summaries, topic tags, and character relationships for easier knowledge management.
    • Structured output: Provides timeline and event list in JSON format for convenient post-processing or knowledge base construction.

Enhanced video content understanding (Doubao series)

  • Online
  • Video content understanding operator that supports using the Doubao model to understand video files, including parsing video content and generating natural language descriptions.
  • Compress the video to within 50 MB, and then use the Doubao model for video understanding.
  • Supported video formats: mp4, wmv, webm, mkv, m4v, flv, avi, mov. Due to the variety of video file format variants, not all files can be guaranteed to be recognized. Please test to verify that the file can be recognized correctly.

Video editing

Operator name

Usage

Operator overview

Intelligent video editing

  • Online
  • The intelligent video editing operator leverages a multimodal large model to provide intelligent video editing capabilities, helping users quickly extract valuable content segments from long videos. It supports understanding editing requirements described in natural language, reference image-assisted recognition (such as roles, objects, and scenes), and multi-dimensional video content analysis (visual, subtitles, and plot). It outputs standardized editing decision information, including timestamps, descriptions, tags, and more.
  • Key features:
    • Supports multiple editing scenarios, including role segment extraction, highlight detection, product segment detection, custom editing, and more.
    • Flexible understanding of editing requirements based on natural language descriptions, supporting custom user requirements.
    • Supports reference image-assisted recognition (including characters, objects, scenes, and more).
    • Multidimensional video content analysis (visuals, subtitles, plot).
    • Supports ASR-enhanced semantic understanding, suitable for videos rich in dialogue and without subtitles, improving the smoothness of segment boundaries.
    • Supports rendering of the three key elements of short dramas (title, prompt text, corner tag), suitable for vertical short drama scenarios.
    • Supports highlight intro functionality, automatically extracting a 10–15 second attractive segment as the opening.
    • Standardized editing decision output (timestamps, descriptions, tags, and more).
    • Automatically generates video segment files and uploads them to TOS.

Viral content editing

  • Online
  • The viral content editing operator can automatically perform shot detection, semantic analysis, editing decisions, and video synthesis based on multiple short drama episodes. It generates multiple promotional videos suitable for traffic acquisition on short video platforms in batches.
  • Key features:
    • Batch processing of multiple short drama episodes: Input multiple short drama episodes at once, automatically complete all analysis and editing, and output multiple viral content in batches.
    • Intelligent editing plan generation: Automatically analyze plot content and generate various editing plans (sequential cut/jump cut) to meet different distribution needs.
    • Automatic cleaning of non-primary content: Automatically detect and remove irrelevant frames such as freeze frames and repeated content from previous or subsequent episodes.
    • Shot-level semantic understanding: Perform plot understanding and importance rating for each shot, outputting structured shot analysis results.
    • Script reconstruction: Automatically reconstruct the complete script of multiple short drama episodes, including character relationships, scene structure, and plot context.
    • Rendering of the three key elements of short dramas: Supports adding drama title, prompt text, and corner tags to the material, adapting to vertical short drama scenarios.
    • Highlight intro: Supports placing highlight segments at the beginning of the material to quickly attract viewer attention.

Video translation

Operator name

Usage

Operator overview

Video translation

  • Online
  • Video translation operator efficiently and accurately converts video content from the source language to one or more target languages. The service covers not only subtitle translation but also speech translation, ultimately outputting the dubbed video and subtitle files in the corresponding language.
  • Key features:
    • Multilingual support: Supports translation in multiple languages. Input language supports video input in 25 languages, and output language supports audio dubbing in 31 languages, covering common languages such as Chinese, English, Japanese, Indonesian, Spanish, Portuguese, Korean, French, and German. Leveraging the powerful translation capabilities of large models, achieves extremely high translation accuracy and terminology localization, meeting the needs of global content dissemination.
    • Voice cloning: Accurately extracts the speaker's voice from the video, achieving a 1:1 restoration of the speaker's vocal characteristics. Meanwhile, the translated audio can be precisely aligned with the original video duration, ensuring the smoothness and consistency of the video.
    • Convenient result delivery: Supports securely writing translated subtitles, dubbed voice audio, and video directly to your specified cloud storage (TOS), while also generating a pre-signed download link.

Video subtitle translation

  • Online
  • Video subtitle translation operator that supports extracting subtitles from videos and translating them into multiple languages. Users can choose to recognize embedded subtitles in the video using OCR, or extract audio subtitles using ASR, then refine and translate the recognized subtitles, and output subtitle files in multiple formats.
  • Key features:
    • Dual subtitle sources: Supports both OCR recognition of embedded subtitles and ASR extraction of audio subtitles.
    • Multiple accuracy levels: Each subtitle source supports both low and high accuracy configurations to meet different scenario requirements.
    • Multilingual translation: Supports subtitle translation in 25 languages, including Chinese, English, Japanese, Korean, French, German, Spanish, and more.
    • Multiple format output: Supports simultaneous output of subtitle files in multiple formats.
    • Bilingual subtitle recognition: OCR mode supports recognition of both translation-type bilingual subtitles and dialogue-type bilingual subtitles.

Video restoration

Operator name

Usage

Operator overview

Video restoration

  • Online
  • Video intelligent restoration operator, based on multimodal large models, enables intelligent removal of video watermarks and subtitles. Supports automatic detection and removal of unwanted content such as watermarks, subtitles, and scrolling subtitles in videos, outputting the restored video file.
  • Key features:
    • Supports removal of multiple targets: watermarks, subtitles, scrolling subtitles, and more.
    • Intelligent detection based on multimodal large models, accurately locates areas that need restoration.
    • Supports precise mask generation, preserving edge details.
    • Supports segmented video processing for more stable handling of long videos.
    • Automatically handles audio retention without additional operations.
    • Supports outputting TOS addresses, results are automatically uploaded.

Subtitle erase

  • Online
  • The subtitle erasure operator automatically detects and erases embedded hard subtitles in the video frames, outputting the video file after erasure. It is suitable for subtitle erasure in vertical-screen videos with white subtitles, providing an efficient and economical subtitle erasure solution, applicable to general scenarios where cost and speed are important.
  • Core capabilities:
    • Intelligent automatic detection and erasure of embedded subtitle regions within the frame.
    • Optimized for vertical screen white subtitle scenarios, featuring high processing efficiency and low cost.
    • Asynchronous task processing; results are retrieved via polling after submission.
    • Outputs the video file and video duration after erasure.

Subtitle erase Pro

  • Online
  • The refined subtitle erasure operator automatically detects and removes embedded hard subtitles from the video frame, outputting the video file after erasure. Suitable for subtitle erasure in videos with vertical screen white subtitles. Compared to the standard version, it delivers superior erasure results without noticeable blurring, accurately reconstructs background textures, and restores the original video frame to a greater extent. It is ideal for scenarios with extremely high image quality requirements, such as the overseas release of short dramas and professional derivative works.
  • Key features
    • Intelligent automatic detection and erasure of embedded subtitle regions within the frame.
    • Supports specified region erasure: the region to be erased can be precisely specified using proportional coordinates.
    • High-quality, seamless erasure with more complete detail retention, no obvious blurring, and higher image quality.
    • Supports Subtitle / Text erasure modes, covering subtitle and on-screen text scenarios.
    • Supports Quality / Size output encoding strategies, balancing image quality and file size.
    • Asynchronous task processing; results are retrieved via polling after submission.

Video processing

Operator name

Usage

Operator overview

Video resolution adjustment (online)

  • Online

Video resolution adjustment operator, key features:

  • Intelligently adjust the resolution to the specified range
  • Supports multiple aspect ratio preservation strategies
  • Video quality and encoding parameters can be controlled
  • Audio stream remains unaffected

Audio and video merging

  • Online
  • The audio and video merging operator uses FFmpeg to sequentially concatenate, adjust the duration of, and finally synthesize the input video and audio materials. The operator supports various input combinations, including one-to-one, one-to-many, many-to-one, and many-to-many. When the total durations of the video and audio are inconsistent, it will automatically select speed alignment or trim to the shorter duration based on the configuration, and upload the resulting video and processing mapping file to TOS.
  • Key features:
    • Supports sequential concatenation of multiple video segments.
    • Supports sequential concatenation of multiple audio segments.
    • Supports preprocessing video and audio separately to the target duration before merging.
    • Supports automatic selection of alignment strategy: prioritizes speed adjustment, and automatically trims when exceeding the threshold.
    • Supports output of the final video file and mapping file, making it easy to track the input, duration, and alignment strategy for each merge.
    • The output directory is automatically isolated by account, request chain, and input hash to prevent results of different tasks from overwriting each other.

Video frame sampling

  • Online
  • The video frame extraction operator supports extracting frames from the input video at a specified frame rate and uploading the extracted image frames to the specified TOS storage path. After the task is completed, the video metadata and the access address for each frame can be obtained. The operator supports configurable frame extraction frequency, maximum frame count limit, output image format, and scaling strategy, making it suitable for scenarios such as video understanding, detection, review, summarization, and cover generation.
  • Key features:
    • Frame extraction by frame rate: Uniformly extracts frames from the video at the specified FPS (0.1 ~ 5.0).
    • Maximum frame count limit: The max_frames parameter can be set to control the upper limit of output frames, preventing excessive frames from long videos.
    • Multi-format output: Supports output in both jpg and png image formats.
    • Flexible scaling: Supports scaling by short side (resize_short_side) or specifying the target resolution (resize_hw) to meet different scenario requirements.
    • Custom output path: Customize the storage path of extracted frame images on TOS using an output path template.
    • Long video support: Suitable for processing long-duration videos; you can control the output scale using max_frames.

Video super-resolution

  • Online
  • The video super-resolution online service enhances the clarity and resolution of input videos based on a video super-resolution model, outputting higher-resolution video results. Applicable to scenarios such as restoration of old films, content enhancement, 4K production, and video sharpening.
  • Key features
    • Supports specifying the target resolution via target_width.
    • Supports automatically maintaining video orientation and inferring the target resolution.
    • Supports automatically retaining the original video audio and uploading it to TOS.

Video frame interpolation

  • Online
  • The video frame interpolation operator is used to increase the frame rate of input videos by generating intermediate frames, thereby improving playback smoothness and outputting new high-frame-rate video files. You can specify the target frame rate, select the interpolation mode as needed, and decide whether to retain the original audio stream.
  • Limitations
    • The input video must be accessible by the service and supports http/https and tos://.
    • output_tos_path must be a TOS directory writable by the current account.
    • target_fps must be greater than 0 and cannot be lower than the source video frame rate.
    • The current version limits video duration to 3 hours and video size to 10GB.
    • The operator depends on CUDA GPU and a video frame interpolation model; the higher the resolution, target frame rate, and video length, the longer the overall processing time.

Face blur

  • Online
  • Face blurring operator, an automated face blurring tool for video content. The operator can automatically detect faces in videos and apply blurring to faces based on the blur level specified by the user, protecting user privacy.
  • Key features
    • Automatically detects faces in video frames and applies blurring
    • Supports multiple blur types (mosaic, Gaussian)
    • Supports fine-grained blurring for different regions (such as elliptical face blurring, face-fitting blurring, eye region blurring, and more)
    • Outputs a unified path for the blurred video (the video will be re-encoded and output even if no faces are detected)

Audio

Audio recognition

Operator name

Usage

Operator overview

Speech-to-text (Doubao Speech ASR)

  • Online
  • The speech-to-text (Doubao series) operator is a speech recognition module that provides a recording transcription solution based on the LAS ASR service.
  • Key features
    • Connects to the Volcano Engine LAS ASR interface
    • Supports automatic sentence segmentation, number normalization, and optional speaker or channel separation
    • Processes multiple audio files concurrently, providing both structured JSON and readable text outputs
  • Suitable for transcribing audio files up to 2 hours in length, supporting advanced features such as punctuation restoration, automatic sentence segmentation, and speaker diarization.

Speech-to-text (Doubao–recording file recognition) enhanced version

  • Online
  • The enhanced LAS speech-to-text (Doubao audio file recognition) operator is based on the Doubao audio file recognition large model and can transcribe speech from input audio or video files into text output. Supports multiple audio and video formats, multiple languages, audio noise reduction, and large file processing. Applicable to scenarios such as content quality inspection and review, audio and video subtitle generation, voice search, and classroom content analysis.
  • Key features
    • Multi-format audio and video input recognition:
      • In addition to audio, video file input is now supported. The LAS operator can automatically extract the audio track from video files for recognition.
      • In addition to raw, wav, mp3, and ogg, support is extended to container formats such as mp4, mov, mkv, and flac.
      • The LAS operator imposes no file size or duration limits on input audio and video files.
      • In addition to public https URL access, TOS internal path access (tos://bucket-name/path/filename) is also supported.
    • Enhanced audio pre-processing to improve model performance:
      • The built-in audio noise reduction module effectively reduces the impact of background noise on recognition, improving the accuracy of audio file transcription.
    • Multi-language support:
      • Can automatically detect the language or recognize speech in a user-specified language.
      • Recognition language support has been expanded to 99 languages, meeting the needs of multi-language and multi-region audio data processing.

Speech-to-text (Doubao-Seed-2.0-lite) enhanced version

  • Online
  • The enhanced LAS speech-to-text (Doubao-Seed-2.0-lite) operator, based on underlying SeedASR recognition, enables Doubao-Seed-2.0-lite multimodal refinement by default: it first completes input download, audio extraction, duration detection, and basic ASR recognition, then sends audio slices, recognition drafts, and hotword context to the multimodal model for refinement, correcting proper nouns, numbers, punctuation, semantic coherence, and supplementing missed content. Supports multiple audio/video formats, multiple languages, speaker separation, hotword enhancement, and sensitive word protection. Suitable for scenarios with high text quality requirements, such as meetings and interviews, podcast courses, and content quality inspection.
  • Key features:
    • Multimodal refinement (enabled by default):
      • Based on the initial recognition results, the Doubao-Seed-2.0-lite multimodal model is invoked to correct and supplement proper nouns, numbers, punctuation, and semantic coherence, significantly improving the quality of the transcribed text.
      • If refinement fails, the system can revert to the basic ASR results to ensure usability.
    • Multi-format audio and video input recognition:
      • Supports audio and video file input, automatically extracting the audio track from video files for recognition.
    • Comprehensive recognition capabilities:
      • Supports text normalization (ITN), automatic punctuation, semantic smoothing (DDC), and speaker separation.
    • Hotwords and sensitive words:
      • Supports direct transmission of hotwords and hotword list enhancement, with hotwords incorporated into both the underlying ASR and refinement prompt.
      • Supports sensitive word protection: system-desensitized content can be excluded from the refinement model, and custom blanking or asterisk replacement is supported.

Audio processing

Operator name

Usage

Operator overview

Audio format conversion (online)

  • Online
  • "Audio format conversion" operator. The audio format conversion operator is used to uniformly convert audio or video files to a specified audio format and output them to a specified storage path.
  • This operator is mainly used in scenarios such as audio format standardization in data processing pipelines, extracting audio from video, and preparing training data. It supports batch concurrent processing and configurable audio encoding parameters.
  • Key features
    • Unified conversion of audio/video to audio
    • Supports custom output audio formats
    • Supports custom output paths (TOS)
    • Supports extension of audio encoding parameters
    • Batch concurrent processing capability

Audio segmentation

  • Online
  • The "audio segmentation" operator is used to extract audio from audio or video files and segment the audio into multiple clips according to specified rules, outputting them to a user-specified storage path.
    This operator is mainly used for structured processing of long audio or video, such as audio preprocessing, data segmentation, and training data construction. It supports batch concurrent processing and flexible output path organization.
  • Key features
    • Audio/video extraction and segmentation
    • Supports custom segmentation rules
    • Supports custom output audio formats
    • Supports output path templates
    • Supports extension of audio encoding parameters

Image

Image generation

Operator name

Usage

Operator overview

Image generation (Seedream series models)

  • Online
  • The image generation (Seedream series models) operator can generate high-quality images based on user input text or reference images, supporting group image generation and streaming output.
  • Key features
    • Text-to-image / image-to-image / group image generation
    • Supports streaming output (SSE) and non-streaming output
    • Output formats supported: URL and b64_json

Image processing

Operator name

Usage

Operator overview

Image resampling

  • Online
  • The image resampling operator is used to resize input images (downsampling only) and save the results to the user-specified TOS directory. Supports four interpolation algorithms (nearest, bilinear, bicubic, lanczos) and .jpg / .png output formats, suitable for image preprocessing, data standardization, offline dataset construction, and more.
  • Key features
    • Multiple interpolation algorithms
      • nearest: Fastest speed, suitable for pixel-style images
      • bilinear: Balanced speed and quality
      • bicubic: Smoother, higher-quality scaling
      • lanczos: Better anti-aliasing, suitable for photos
    • URL / TOS input support
      • image_src_type=image_url: Input public URL
      • image_src_type=image_tos: Input tos:// address
    • Output to TOS
      • tos_dir (required): specifies the output directory (folder level)
      • The output file name is generated by the server (can be assisted by image_name), and the _resample suffix is appended for identification
    • Output format and DPI control
      • Output formats supported .jpg / .png
      • Supports setting output DPI (target_dpi)

Document

Document parsing

Operator name

Usage

Operator overview

PDF document parsing (Doubao)

  • Online
  • The PDF content parsing operator supports visual model parsing of PDF files and structured output in Markdown format.
  • Key features
    • Supports PDF page rendering and visual model parsing, outputs high-fidelity Markdown, and fully restores the original structure (heading hierarchy, tables, formulas, image regions).
    • Automatically identifies image regions and returns bounding box information and pre-signed image URLs.
    • Supports both page-by-page and entire-book Markdown aggregation for subsequent content processing and presentation.

Multimodal

Multimodal vectorization

Operator name

Usage

Operator overview

Image-text embedding (Doubao series models)

  • Online
  • Offline
  • The multimodal vector generation processor supports joint vector generation for images/videos and text, enabling cross-modal retrieval capabilities.
  • Key features
    • Multimodal vectorization support: enables joint vector generation for images/videos and text, providing cross-modal retrieval capabilities.
    • Input format adaptation:
      • Natively supports input formats such as base64 encoding, binary data, and URLs for images/videos
      • Automatically handles media format conversion (JPEG/PNG/MP4/AVI and more)

Multimodal deep reasoning

Operator name

Usage

Operator overview

Multimodal deep reasoning (Doubao-seed-2.0)

  • Online
  • Provides deep reasoning capabilities of large models in multimodal scenarios, analyzing and understanding images, videos, or text, and returning structured text output. The operator automatically constructs a message structure that complies with multimodal model specifications. Users only need to provide image, video, or text data as specified to complete inference.
  • Key features
    • Deep reasoning mechanism: the model automatically decomposes questions and performs logical reasoning before answering, generating a reasoning chain (reasoning_content)
    • Multimodal scenario support: supports simultaneous input of images, videos, and text, and automatically assembles multimodal messages
    • Input simplification mechanism: supports multiple data sources such as local files, HTTP/HTTPS URLs, and TOS/S3 object storage, enabling visual understanding capabilities through simple configuration
    • Flexible thinking mode: Supports controlling the deep thinking mode via the thinking_type parameter (enabled / disabled / auto), allowing flexible trade-offs between answer quality and performance.

Multimodal deep reasoning (Doubao-seed-1.8)

  • Online

Vector retrieval

General retrieval

Operator name

Usage

Operator overview

Lance General Search

  • Online
  • The Lance general retrieval operator provides unified and efficient retrieval capabilities for Lance and LanceDB format data stored in Volcano Engine Object Storage (TOS), supporting vector retrieval, full-text retrieval, hybrid retrieval, and scalar queries. It is suitable for large-scale structured and unstructured data retrieval scenarios.
  • Key features:
    • Multi-mode retrieval support: Supports four retrieval modes—vector retrieval (vector), full-text retrieval (fts), hybrid retrieval (hybrid), and scalar query (scalar).
    • Multimodal input: Vector retrieval supports three input types—vector arrays, text strings, and Base64-encoded images.
    • Data filtering: Supports data filtering using SQL WHERE clauses for pre-filter or post-filter operations.
    • Result re-ranking: The hybrid retrieval mode supports the RRF (Reciprocal Rank Fusion) re-ranking algorithm.
    • Column selection: Supports specifying returned columns to reduce data transmission.
Last updated: 2026.08.06 09:58:41