TrueSeeker AI · Verified claim report Case f83779d06d · 2026-08-24

§ Claim under review · Capability

"LongCat-Video is an open-source model from Meituan that can turn a single photo and an audio track into a talking video with synchronized facial and mouth movements. The project supports more than realistic human faces. The same system can animate anime characters, animals, and other stylized subjects, while handling both single-speaker and multi-speaker audio. For speech alignment, the pipeline uses Whisper-Large-v3 to help extract audio features and keep lip movements matched with the spoken words, including across longer generated clips. Because the code and model are publicly available, developers can run, test, and build on the technology themselves instead of relying entirely on closed commercial avatar platforms." (Source given: https://github.com/meituan-longcat/LongCat-Video)

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

Medium
§

Summary

This post is mostly right but it names the wrong model. Meituan has indeed open-sourced a system that turns one photo plus an audio track into a lip-synced talking video, and it does support anime characters, animals, and conversations with two audio streams. However, that system is called LongCat-Video-Avatar 1.5, released in May 2026. LongCat-Video, the name used in the post, is the underlying 13.6 billion parameter base model, and on its own it only does text-to-video, image-to-video, and video continuation with no audio input at all. The Whisper-Large-v3 detail is also version-specific: it applies to version 1.5, while the earlier December 2025 avatar release used a different audio encoder called wav2vec2. The GitHub link in the post is correct, since the avatar code lives in that same repository, and the licence really is MIT, so anyone can run and build on it. One caveat worth keeping in mind is that all the quality claims, including the anime and animal support and the stability over long clips, come from Meituan's own model card and its own unreviewed technical report, with no independent testing found.

§

The readings

key figures from the evidence
13.6 B parameters

LongCat-Video base foundation model size

508 image-audio pairs

vendor benchmark size for v1.5 quality claims

8 steps

inference acceleration via step distillation in v1.5

§

Why this verdict

Every substantive capability in the post checks out against Meituan's own primary artifacts, and the linked repo is the correct one, but the model is misnamed: the audio-driven system is LongCat-Video-Avatar-1.5, released 2026-05-21, not the LongCat-Video base model released 2025-10-25, which does text-to-video, image-to-video and video-continuation only. I considered "Partially accurate but misleading" because the name error is the kind that sends a reader to the wrong weights, and rejected it because the operative proposition, that Meituan has open-sourced a photo-plus-audio talking-avatar system with anime, animal and multi-speaker support built on Whisper-Large-v3, is fully supported and the cited repo does contain that code. I considered "Source exists but framing is misleading" and rejected it for the same reason: nothing here is exaggerated in kind, only imprecise in naming and version scope. Confidence is capped at Medium rather than High because the claim is version-ambiguous by the rule that you cannot be highly confident about a subject the claim did not pin, and because all capability descriptions rest on vendor self-report with no independent evaluation found. Assessment is as of 2026-08-24. ---
§

Evidence

There are two distinct things in this family, and the claim merges them.

The base model. LongCat-Video is a foundational video generation model with 13.6B parameters covering Text-to-Video, Image-to-Video and Video-Continuation, and it unifies those three tasks in a single framework . It was released Oct 25, 2025 . Audio conditioning is not among its listed tasks; its Hugging Face card tags it text-to-video, image-to-video and video-continuation .

The avatar model. The audio-driven system is a separate checkpoint line. LongCat-Video-Avatar was announced Dec 16, 2025 as a unified model for audio-driven character animation supporting Audio-Text-to-Video, Audio-Text-Image-to-Video and Video Continuation, with compatibility for single-stream and multi-stream audio inputs . On May 21, 2026 Meituan released LongCat-Video-Avatar-1.5, which replaces Wav2Vec2 with Whisper-Large for more accurate lip synchronization, generalizes to stylized domains (anime, animals, complex real-world conditions), supports single-stream and multi-stream audio, and accelerates inference to 8 steps via step distillation . Its card states it is built on the LongCat-Video foundation model and supports AT2V, ATI2V and Video Continuation, and lists stylized domain generalization to anime, animals and multi-person interactions .

The encoder detail is version-gated. The card states that --model_type avatar-v1.0 uses the wav2vec2 audio encoder by default, while --model_type avatar-v1.5 uses the Whisper-large-v3 audio encoder for better lip sync quality .

The avatar code ships inside the repo the post links to: LongCat-Video-Avatar extends the foundational LongCat-Video pipeline by adding audio conditioning, reusing the core model components while introducing audio processing modules, and supports single-audio input for one character and multi-audio input with two streams for multi-character dialogues . A third party notes the dependency explicitly: it depends on two weight directories, LongCat-Video as the base video generation model and LongCat-Video-Avatar-1.5 as the avatar model .

Licensing: the LongCat-Video model weights are released under the MIT License , and a practitioner review describes 1.5 as an open-weight audio-driven human video model built on LongCat-Video and released under MIT .

Quality claims are vendor-run. The tech report states v1.5 achieves competitive or superior performance versus leading closed-source systems such as HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 on human-likeness and expert quality assessments on the authors' own benchmark , a benchmark of 6 application scenarios, 2 languages and 2 visual styles totalling 508 image-audio pairs, scored by 770 crowdsourced evaluators producing 13,240 judgments plus 10 domain experts .


§

Findings

✓ What's accurate 6

  • Meituan's LongCat team has publicly released an open-source, MIT-licensed audio-driven avatar system that turns a reference image plus audio into a lip-synced talking video (Audio-Text-Image-to-Video).
  • It supports stylized subjects: the card lists robust generalization to anime, animals and complex real-world conditions such as multi-person interactions.
  • It supports both single-stream and multi-stream audio inputs, with documented merge and concatenation modes for two speakers.
  • The v1.5 checkpoint does use a Whisper-large-v3 audio encoder, replacing wav2vec2, for better lip sync.
  • Long-clip handling is a stated design goal: the report describes accurate lip-synchronization, full-body temporal stability and robust long-video generation with strict identity consistency.
  • Code and weights are publicly downloadable from the GitHub repo the post cites, so the "developers can run and build on it" framing is correct.

≈ What's misleading 4

  • **Scale conflation:** the post attributes all of this to "LongCat-Video." That name belongs to the 13.6B foundation model whose documented tasks are Text-to-Video, Image-to-Video and Video-Continuation, with no audio conditioning. The audio-driven behaviour belongs to a separate checkpoint line, LongCat-Video-Avatar and LongCat-Video-Avatar-1.5. A reader who downloads the weights named in the post gets a model that cannot do what the post describes; they need both weight directories. The mitigating fact is that the avatar code lives in the linked repo, so the URL is right even though the model name is not.
  • **Omitted qualifier:** "the pipeline uses Whisper-Large-v3" is true only of v1.5. v1.0, the default model_type, uses wav2vec2. Stated without the version, it describes the family incorrectly for the checkpoint released in December 2025.
  • **Marketing as evidence:** every capability statement in the post traces to Meituan's own model card and its own unrefereed technical report. The anime and animal generalization, the long-clip stability and the lip-sync accuracy are vendor self-descriptions. That is fine as evidence that the features are offered; it is not independent evidence that they work well.
  • **Capability extrapolation (mild):** "keep lip movements matched with the spoken words, including across longer generated clips" is stated flatly. The underlying source frames long-video stability as an engineering goal validated on the authors' own 508-case benchmark, not as an unconditional property.

? What's uncertain 4

  • How well the stylized-subject and long-clip claims hold under ordinary user conditions. No independent quantitative evaluation of LongCat-Video-Avatar-1.5's lip sync or identity consistency was located, only vendor benchmarks and informal hands-on reports.
  • Whether the post intended v1.0 or v1.5. It is version-ambiguous, and the Whisper detail only fits v1.5, so I resolved it to v1.5.
  • The reported parity or superiority against HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 is entirely vendor-run on a vendor-designed benchmark and is not independently verified. The post does not make this claim, but it is the implicit backdrop of the "instead of closed commercial platforms" line.
  • Exact release date of v1.0 varies between sources (the repo news log says Dec 16, 2025; at least one third-party page says November 2025). I used the repo.
Distortion flags scale conflation omitted qualifier marketing as evidence capability extrapolation
§

Sources

8 of 8 linked to records
[1]

Hugging Face model card, meituan-longcat/LongCat-Video-Avatar-1.5 (vendor artifact of record for the avatar model)

primary vendor
https://huggingface.co/meituan-longcat/LongCat-Video-Avatar-1.5 ↗
[2]

GitHub README, meituan-longcat/LongCat-Video (the repo cited by the post; also hosts the avatar code and release log)

primary vendor
https://github.com/meituan-longcat/LongCat-Video ↗
[3]

LongCat-Video-Avatar 1.5 Technical Report, arXiv 2605.26486

primary vendor lab
https://arxiv.org/abs/2605.26486 ↗
[4]

Hugging Face model card, meituan-longcat/LongCat-Video (base foundation model)

primary vendor
https://huggingface.co/meituan-longcat/LongCat-Video ↗
[5]

LongCat-Video Technical Report, arXiv 2510.22200

primary vendor lab
https://arxiv.org/abs/2510.22200 ↗
[6]

Hugging Face model card, meituan-longcat/LongCat-Video-Avatar (v1.0)

primary vendor
https://huggingface.co/meituan-longcat/LongCat-Video-Avatar ↗
[7]

DeepWiki structural walkthrough of the avatar pipeline in the repo

secondary third-party documentation aggregator
https://deepwiki.com/meituan-longcat/LongCat-Video/5-longcat-video-avatar ↗
[8]

Third-party ComfyUI integration (smthemex/ComfyUI_LongCat_Avatar) and a hands-on review (VisionStory, July 2026)

secondary community implementation / practitioner review
https://github.com/smthemex/ComfyUI_LongCat_Avatar ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →