← October 9, 2026 briefing · added in the 04:00 KST update

Tencent releases 'Youtu-Parsing-Omni', a 5B model that parses documents, images, audio and video into a single JSON (community report)

What happened

According to a Hugging Face model card shared on Reddit, Tencent released Youtu-Parsing-Omni, a 5B omni-modal parsing model. It takes document pages, natural images, charts and flowcharts, geometric figures, audio clips and video as input and outputs a single structured JSON. The output combines perception results, such as layout, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events and camera motion, with understanding results such as captions, descriptions and reports. Task prompts select which results to produce.

Why it matters

Pipelines that used to run document OCR, ASR and video analysis separately could be merged into one small model, lowering RAG and data preprocessing costs.

Confidence medium · official source pending

Sources

More about Tencent