§ Claim under review · Capability
"LongCat-Video is an open-source model from Meituan that can turn a single photo and an audio track into a talking video with synchronized facial and mouth movements. The project supports more than realistic human faces. The same system can animate anime characters, animals, and other stylized subjects, while handling both single-speaker and multi-speaker audio. For speech alignment, the pipeline uses Whisper-Large-v3 to help extract audio features and keep lip movements matched with the spoken words, including across longer generated clips. Because the code and model are publicly available, developers can run, test, and build on the technology themselves instead of relying entirely on closed commercial avatar platforms." (Source given: https://github.com/meituan-longcat/LongCat-Video)
Verdict
Mostly accurate
Confidence
MediumSummary
This post is mostly right but it names the wrong model. Meituan has indeed open-sourced a system that turns one photo plus an audio track into a lip-synced talking video, and it does support anime characters, animals, and conversations with two audio streams. However, that system is called LongCat-Video-Avatar 1.5, released in May 2026. LongCat-Video, the name used in the post, is the underlying 13.6 billion parameter base model, and on its own it only does text-to-video, image-to-video, and video continuation with no audio input at all. The Whisper-Large-v3 detail is also version-specific: it applies to version 1.5, while the earlier December 2025 avatar release used a different audio encoder called wav2vec2. The GitHub link in the post is correct, since the avatar code lives in that same repository, and the licence really is MIT, so anyone can run and build on it. One caveat worth keeping in mind is that all the quality claims, including the anime and animal support and the stability over long clips, come from Meituan's own model card and its own unreviewed technical report, with no independent testing found.
The readings
key figures from the evidenceLongCat-Video base foundation model size
vendor benchmark size for v1.5 quality claims
inference acceleration via step distillation in v1.5
Why this verdict
Evidence
There are two distinct things in this family, and the claim merges them.
The base model. LongCat-Video is a foundational video generation model with 13.6B parameters covering Text-to-Video, Image-to-Video and Video-Continuation, and it unifies those three tasks in a single framework . It was released Oct 25, 2025 . Audio conditioning is not among its listed tasks; its Hugging Face card tags it text-to-video, image-to-video and video-continuation .
The avatar model. The audio-driven system is a separate checkpoint line. LongCat-Video-Avatar was announced Dec 16, 2025 as a unified model for audio-driven character animation supporting Audio-Text-to-Video, Audio-Text-Image-to-Video and Video Continuation, with compatibility for single-stream and multi-stream audio inputs . On May 21, 2026 Meituan released LongCat-Video-Avatar-1.5, which replaces Wav2Vec2 with Whisper-Large for more accurate lip synchronization, generalizes to stylized domains (anime, animals, complex real-world conditions), supports single-stream and multi-stream audio, and accelerates inference to 8 steps via step distillation . Its card states it is built on the LongCat-Video foundation model and supports AT2V, ATI2V and Video Continuation, and lists stylized domain generalization to anime, animals and multi-person interactions .
The encoder detail is version-gated. The card states that --model_type avatar-v1.0 uses the wav2vec2 audio encoder by default, while --model_type avatar-v1.5 uses the Whisper-large-v3 audio encoder for better lip sync quality .
The avatar code ships inside the repo the post links to: LongCat-Video-Avatar extends the foundational LongCat-Video pipeline by adding audio conditioning, reusing the core model components while introducing audio processing modules, and supports single-audio input for one character and multi-audio input with two streams for multi-character dialogues . A third party notes the dependency explicitly: it depends on two weight directories, LongCat-Video as the base video generation model and LongCat-Video-Avatar-1.5 as the avatar model .
Licensing: the LongCat-Video model weights are released under the MIT License , and a practitioner review describes 1.5 as an open-weight audio-driven human video model built on LongCat-Video and released under MIT .
Quality claims are vendor-run. The tech report states v1.5 achieves competitive or superior performance versus leading closed-source systems such as HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 on human-likeness and expert quality assessments on the authors' own benchmark , a benchmark of 6 application scenarios, 2 languages and 2 visual styles totalling 508 image-audio pairs, scored by 770 crowdsourced evaluators producing 13,240 judgments plus 10 domain experts .
Findings
✓ What's accurate 6
- Meituan's LongCat team has publicly released an open-source, MIT-licensed audio-driven avatar system that turns a reference image plus audio into a lip-synced talking video (Audio-Text-Image-to-Video).
- It supports stylized subjects: the card lists robust generalization to anime, animals and complex real-world conditions such as multi-person interactions.
- It supports both single-stream and multi-stream audio inputs, with documented merge and concatenation modes for two speakers.
- The v1.5 checkpoint does use a Whisper-large-v3 audio encoder, replacing wav2vec2, for better lip sync.
- Long-clip handling is a stated design goal: the report describes accurate lip-synchronization, full-body temporal stability and robust long-video generation with strict identity consistency.
- Code and weights are publicly downloadable from the GitHub repo the post cites, so the "developers can run and build on it" framing is correct.
≈ What's misleading 4
- **Scale conflation:** the post attributes all of this to "LongCat-Video." That name belongs to the 13.6B foundation model whose documented tasks are Text-to-Video, Image-to-Video and Video-Continuation, with no audio conditioning. The audio-driven behaviour belongs to a separate checkpoint line, LongCat-Video-Avatar and LongCat-Video-Avatar-1.5. A reader who downloads the weights named in the post gets a model that cannot do what the post describes; they need both weight directories. The mitigating fact is that the avatar code lives in the linked repo, so the URL is right even though the model name is not.
- **Omitted qualifier:** "the pipeline uses Whisper-Large-v3" is true only of v1.5. v1.0, the default model_type, uses wav2vec2. Stated without the version, it describes the family incorrectly for the checkpoint released in December 2025.
- **Marketing as evidence:** every capability statement in the post traces to Meituan's own model card and its own unrefereed technical report. The anime and animal generalization, the long-clip stability and the lip-sync accuracy are vendor self-descriptions. That is fine as evidence that the features are offered; it is not independent evidence that they work well.
- **Capability extrapolation (mild):** "keep lip movements matched with the spoken words, including across longer generated clips" is stated flatly. The underlying source frames long-video stability as an engineering goal validated on the authors' own 508-case benchmark, not as an unconditional property.
? What's uncertain 4
- How well the stylized-subject and long-clip claims hold under ordinary user conditions. No independent quantitative evaluation of LongCat-Video-Avatar-1.5's lip sync or identity consistency was located, only vendor benchmarks and informal hands-on reports.
- Whether the post intended v1.0 or v1.5. It is version-ambiguous, and the Whisper detail only fits v1.5, so I resolved it to v1.5.
- The reported parity or superiority against HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 is entirely vendor-run on a vendor-designed benchmark and is not independently verified. The post does not make this claim, but it is the implicit backdrop of the "instead of closed commercial platforms" line.
- Exact release date of v1.0 varies between sources (the repo news log says Dec 16, 2025; at least one third-party page says November 2025). I used the repo.
Sources
8 of 8 linked to recordsHugging Face model card, meituan-longcat/LongCat-Video-Avatar-1.5 (vendor artifact of record for the avatar model)
GitHub README, meituan-longcat/LongCat-Video (the repo cited by the post; also hosts the avatar code and release log)
LongCat-Video-Avatar 1.5 Technical Report, arXiv 2605.26486
Hugging Face model card, meituan-longcat/LongCat-Video (base foundation model)
LongCat-Video Technical Report, arXiv 2510.22200
Hugging Face model card, meituan-longcat/LongCat-Video-Avatar (v1.0)
DeepWiki structural walkthrough of the avatar pipeline in the repo
Third-party ComfyUI integration (smthemex/ComfyUI_LongCat_Avatar) and a hands-on review (VisionStory, July 2026)