Russian Mathematicians Educate AI Models to Communicate Non-Verbally.

Recently, I had the opportunity to meet some exceptionally talented Russian mathematicians who revealed to me a method for enabling artificial intelligence models to communicate through a mechanism similar to machine telepathy.
These mathematicians are part of a startup named Mostik, which translates to bridge in Russian. This reflects the group’s methodology, allowing various models to interact by using the mathematical values inherent in their weights—elements that influence how a prompt is converted into an output. Essentially, this approach enables the capabilities of a larger model to enhance a smaller model’s intelligence more effectively.
They utilized this technique to develop a model that soared to the top of ARC-AGI 3, a notoriously challenging competition for AI models. (The team remained tight-lipped about further details, as they aim to win the contest.) To illustrate their concept, they created a connection between two Chinese open-weight models: the massive GLM-5.2, equipped with 753 billion parameters, and a mobile-friendly version of Qwen-3.5 with 4 billion parameters. The resulting hybrid system costs just one-twentieth of the full GLM model, achieving performance that is precisely in between the two.
“It’s widely recognized in machine learning that ensembles of models outperform singular ones,” Sasha Malysheva, CEO of Mostik, shared with me over coffee.
Malysheva, who pioneered the approach, humorously recounted an inside joke at the company: predicting the future of AI resembles estimating the weight of a pig. In mathematical circles, it’s a known fact that a group of random individuals can more accurately guess a pig’s weight than a single expert, as their combined estimates yield a better average.
Similar to collaboratively estimating an animal’s weight, merging the outputs of multiple AI models typically yields superior results. Normally, this involves inputting one model’s output into another, which can be time-consuming and costly. However, the Mostik team discovered a method for AI models to communicate directly without generating text output. If successful, this could significantly enhance the value of open-weight models, positioning them to compete more effectively against the proprietary models offered by elite labs like Anthropic and OpenAI.
Malysheva believes that conjoining various models may prove to be a more effective path for advancing AI. “I personally doubt that we will see a singular monolithic model [in the future] or that the power of models will solely derive from scale,” she remarked, referencing the trend of increasing model sizes and feeding them vast datasets.
“If Mostik can facilitate pairing frontier models with domain-specific ones—like those in biology or physics—this would lead to the training of many more specialized models,” explained Vladimir Arustamian, the tech lead at AI software firm Lovable, who is familiar with the Mostik team. “This team has been working on this for just a few months, and they’ve already achieved something I would have predicted would take years.”
The Mostik method allows “you to approach large-model quality without one large model managing the entire process, providing considerable improvements with just a smaller model operating in tandem,” stated Karl Tuyls, a former computer scientist at Google DeepMind acquainted with the company’s technology. He added that the method is an obvious choice for those aiming to run models as efficiently as possible.
Stanislav Smirnov, a professor at the University of Geneva and a Fields Medal recipient in 2010, serves as Mostik’s chief scientist. He expressed that finding common ground between two AI models presents unexpected challenges. “There seems to be no suitable mathematical language yet,” he noted. In the meantime, Mostik’s approach literally bridges that gap.
