AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Thinking Machines has introduced Inkling, a 975-billion-parameter multimodal AI model, available on Hugging Face. Its size and open access could transform AI applications, but hardware requirements and evaluation details remain uncertain.

Thinking Machines has released Inkling on Hugging Face, a 975-billion-parameter multimodal model capable of processing text, images, and audio within a one-million-token context window. For more details, see Could Thinking Machines’ Inkling Be The Key To AI’s Next Leap?. The release makes a large-scale, open multimodal model accessible, though running it requires substantial hardware resources. This development is discussed in the original analysis Welcome Inkling By Thinking Machines. This development could influence future AI research and application development.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities. Its architecture employs 256 experts, combining global and sliding-window attention, with specialized modules for image and audio inputs. To understand the significance of this architecture, see the detailed analysis here. The model is available via Hugging Face with support for popular inference frameworks, but hardware requirements are extremely high: around 2 TB of VRAM for BF16 checkpoints and approximately 600 GB for NVFP4 versions.

While the model is positioned for domain-specific fine-tuning in scientific, media, and enterprise contexts, independent benchmark results, safety evaluations, and licensing details remain undisclosed. The release includes checkpoints optimized for lower-precision inference, but practical deployment is limited to organizations with powerful hardware or access through hosted services.

At a glance
reportWhen: announced July 2026
The developmentThinking Machines has released Inkling, a large-scale multimodal AI model, on Hugging Face, marking a significant step in open AI development with high hardware demands.

Potential Impact of Large-Scale Multimodal Models

This release signifies a step toward more capable and accessible multimodal AI, which can reason across text, images, and audio within a single framework. If effectively deployed, Inkling could advance applications in scientific research, media analysis, and enterprise workflows. However, the high hardware demands and lack of independent evaluation currently limit immediate broad adoption, raising questions about its practical usability and safety.

Amazon

high VRAM graphics card for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Large Multimodal AI Development

Recent years have seen rapid growth in large language models, with multimodal variants emerging to handle multiple data types simultaneously. Prior models, such as GPT-4 and PaLM-E, have demonstrated multimodal capabilities but often remain proprietary or limited in scale. Inkling’s release on Hugging Face as an open model at this scale is notable, given the historical trend toward more restricted access and the increasing hardware requirements for training and deployment.

Earlier efforts focused on smaller models or specialized systems; Inkling’s scale and multimodal nature mark a significant evolution, although independent benchmarking and safety assessments are still pending. The model’s architecture, based on sparse Mixture-of-Experts design, aims to balance scale with efficiency, but real-world performance remains to be validated.

“This model is huge.”

— Hugging Face spokesperson

Amazon

professional AI inference hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

Details about independent benchmark results, safety evaluations, and licensing terms are not yet available. It is unclear how well Inkling performs across different modalities in real-world tasks or how its speed and accuracy compare to other models. The practical implications of its high hardware requirements and the performance of quantized versions remain to be seen.

Amazon

multimodal AI model server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Evaluation

Developers and organizations are expected to begin testing Inkling through supported inference frameworks. Early evaluations will focus on latency, memory use, and multimodal accuracy. Additionally, independent benchmarking, safety testing, and domain-specific fine-tuning will clarify the model’s capabilities and limitations. The release of detailed model cards, licensing information, and performance benchmarks will be critical in assessing its broader adoption.

Amazon

large-scale AI training GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling and what makes it significant?

Inkling is a large-scale multimodal AI model from Thinking Machines, with 975 billion parameters, capable of understanding text, images, and audio within a single framework. Its open release on Hugging Face marks a notable development in accessible, high-capacity AI models.

Can Inkling process videos or real-time multimedia?

While the architecture supports image inputs with a temporal dimension, native video processing has not been evaluated. Its potential for video tasks remains speculative until further testing confirms its capabilities.

Is Inkling deployable on consumer hardware?

Currently, the hardware requirements are extremely high—around 2 TB of VRAM for BF16 checkpoints—making full deployment impractical for typical consumer systems. Most users will need access via cloud services or specialized hardware.

Will the model be safe and ethically evaluated?

No independent safety or bias evaluations have been published yet. The lack of detailed safety assessments means caution is advised when considering deployment in sensitive applications.

What are the next steps for researchers and developers interested in Inkling?

They should begin testing the model through supported inference engines and monitor upcoming publications of benchmark results, safety evaluations, and licensing details to better understand its practical performance and limitations.

Source: ThorstenMeyerAI.com

You May Also Like

The High-End PC and Workstation Tax

Memory costs surge in 2026, making high-end PC building more expensive and challenging DIY efforts. Prebuilts may now be more cost-effective.

Adam Mosseri Surges In Global Coverage

Adam Mosseri, head of Instagram, experiences a surge in international media coverage, with 22 mentions in recent reporting, marking a notable increase.

EVE Online’s Carbon Engine Is Now Open Source: Fenris Creations Explains Why

Fenris Creations has announced the open-source release of EVE Online’s Carbon engine, explaining their reasons for making the code publicly available.

What PoE Switches Add to Larger Camera Installations

Find out how PoE switches simplify large camera setups and why they’re essential for seamless, scalable security installations.