AI & MLChatGPTLLM

Can AI enhance our ability to communicate complex concepts?

In 2023 I wrote about using multimodal generative AI to approximate the “Gestalt” communication in Greg Egan’s 1997 hard science-fiction novel Diaspora. Developments in generative AI, neuroscience and AI-to-AI communication since then make parts of the idea more plausible.

However, my original post—and some answers from ChatGPT 4—overstated what Egan describes. The question also needs refining: beyond communicating more information, can AI help two people construct sufficiently similar internal models of an idea, while transmitting much less explicit explanation?

Gestalt communication in Diaspora

Diaspora begins in 2975 CE, with humanity divided into fleshers (biological humans and their modified descendants), Gleisner robots (software minds in physical robotic bodies), and polis citizens (software people living in virtual communities called polises and environments called scapes).

Egan describes two standardized information modalities used by citizens: linear and gestalt, distant descendants of hearing and vision.

Gestalt is not telepathy or two minds merging into one consciousness. That was an embellishment in my original interpretation. It appears to be a rich perceptual and representational channel combining:

  • images, symbolic information and non-visual tags;
  • object metadata, such as an asteroid’s chemical composition, mass, spin and orbital parameters;
  • “colour” beyond the human visual spectrum;
  • identity and other information broadcast by citizens.

Think of a standardized multidimensional representation that presents perception, metadata, meaning and context together.

Gestalt is higher-dimensional communication, not unlimited communication.

Can generative AI communicate in Gestalt?

When I asked ChatGPT 4 whether it could communicate like citizens of Diaspora, it said no. Literally, that remains true: humans lack Egan’s standardized Gestalt input channel. A language model cannot send me an arbitrary multidimensional structure and have my brain directly interpret it.

Revisiting the proposal with GPT-6 Astra, the question has changed: could AI help one person communicate a rich internal model to another through a shared representation and personalized explanation? The assessment is a qualified yes to investigating that approximation. This is not evidence that a newer model has demonstrated literal Gestalt communication:

Human A → AI → shared semantic representation → AI → Human B

A supplies conversation, text, diagrams, documents, photographs, equations or other material. An AI structures what A appears to mean; another AI process reconstructs it for B using what B already understands.

The objective would be a sufficiently similar conceptual state, not an identical brain state. The central problem becomes semantic alignment. The receiver should be able to reconstruct the sender’s reasoning without having to adopt it as their own.

Communication as model alignment

When I explain an idea, I turn part of my internal model into words. You reconstruct your own model from them. Communication works when those models become sufficiently similar for our purpose. This relates to Clark and Brennan’s account of grounding in communication: establishing sufficient understanding for the purpose at hand.

The interactive-alignment model of dialogue developed by Martin Pickering and Simon Garrod proposes that people progressively align linguistic representations and develop increasingly compatible situation models.

Research from Uri Hasson’s lab at Princeton demonstrated measurable coupling between speakers’ and listeners’ brain activity, with greater coupling associated with greater understanding.

In a 2024 Neuron paper, the same group used contextual embeddings from large language models to model linguistic information passing from a speaker’s brain, through speech, to a listener’s brain. These embeddings could model a shared, context-rich linguistic space involved in communication.

This does not mean brains operate like GPT or contain literal embedding vectors. It supports the idea that successful communication appears to involve constructing related semantic representations in different minds. Words carry the message; meaning is reconstructed at either end.

The bottleneck may not be information bandwidth

My original post emphasized augmented reality, sound, haptics and brain-computer interfaces. These could increase bandwidth, but bandwidth is probably not the main constraint.

I can download a 500 MB textbook in seconds without understanding it. Ten simultaneous information streams in an augmented-reality headset might make comprehension harder. Attention, working memory, prior knowledge and conceptual integration remain constraints.

The objective should therefore be closer to maximize useful conceptual alignment per unit of human attention than maximize information transmitted per second. Generative AI could help by adapting the explanation to those constraints.

What would a human Gestalt system actually do?

Explaining a complicated idea can take hours or even days of discussion, diagrams, reading, examples and analogies. The listener may need to learn unfamiliar concepts, connect them to experience and work out how they fit—or conflict—with an existing worldview. I might clarify assumptions, uncertainties and where an analogy breaks down, answer questions, then discover we interpreted a term differently. The work includes learning and reflection, not just delivering an explanation.

A Gestalt-like system would capture the structure behind those activities from rough notes, speech, sketches and supporting material, without requiring a polished presentation. Its intermediate representation could contain:

  • concepts, relationships, causal structures and dependencies;
  • assumptions, evidence, confidence, uncertainty and disagreements;
  • examples, counterexamples, metaphors and analogies;
  • visual or spatial representations;
  • relative importance and the sender’s intended conclusion.

Call this a Gestalt packet: a machine-readable representation of the sender’s intended meaning, not necessarily something a person reads directly. An initial packet could simply organize text, relationships and source links; it would not need to transmit a model’s hidden internal states.

The receiving AI becomes a personalized decoder

The receiving AI would ideally know something about your existing conceptual model. Suppose I explain a distributed organizational problem through biology, but you know distributed software systems better. Your AI might reconstruct the explanation using cache coherence, distributed state, synchronization, consensus and propagation delays. Dedre Gentner’s structure-mapping theory of analogy provides useful background here: the connections between things matter more than their surface resemblance.

Another receiver might get an organizational diagram, worked example or story. Someone already familiar with the concepts might need only a compact diagram and three sentences. The underlying representation is similar; the reconstruction is different.

Teachers, writers and people who know us well already perform this translation. AI could potentially do parts of it dynamically for each participant, reducing the sender’s burden. That seems more plausible than sending thoughts directly between brains.

A lo-fi Gestalt system could be built with current technology

Most individual components already exist; combining them would be the experiment.

Multimodal input

Modern generative AI systems can process combinations of text, speech, images, documents and structured data.

Semantic modelling

Language models can identify the packet’s concepts, relationships, assumptions, contradictions and analogies. Knowledge graphs could supplement prose representations.

Personalization

The decoder could use vocabulary, expertise, previous conversations and preferred explanations to adapt its reconstruction.

Multimodal output

Possible renderings include text, diagrams, interactive concept maps, animation, audio, simulations, equations, examples, timelines and spatial interfaces.

Feedback

Instead of merely asking “Did you understand?”, the system could test for mismatches:

  • ask counterfactual questions or request predictions;
  • show two interpretations and ask which the receiver inferred;
  • check whether participants use the same word with different meanings.

Feedback could guide an iterative process of model convergence. Rozenblit and Keil’s research on the illusion of explanatory depth gives one reason to look beyond self-reported understanding.

Who is doing something similar?

Nobody appears to be building the complete system described above, but several fields are exploring its components.

Uri Hasson’s lab — shared linguistic representations between humans

The speaker-listener studies discussed earlier examine the phenomenon this system would try to improve: the transfer and reconstruction of meaning between minds. They investigate human communication rather than implement a Gestalt system.

Interlat — AI agents communicating without language

A paper presented at ACL 2026 introduced Interlat, which lets AI agents communicate through continuous internal latent representations rather than first translating them into natural-language tokens.

This avoids reducing a rich representation to a sequence of discrete tokens before another model processes it. In their experiments, the researchers reported better performance on some tasks and substantial communication compression.

This resembles one part of the Diaspora idea: AI internal representation → shared latent representation → AI internal representation. Humans cannot consume those representations directly. A human Gestalt system would still need AI to translate between the shared representation and each person’s conceptual structures.

BrainNet — literal brain-to-brain communication

In 2019, researchers at the University of Washington and Carnegie Mellon demonstrated BrainNet. It connected three people through EEG and transcranial magnetic stimulation: two communicated decisions to a third during a simplified Tetris-style task.

This demonstrated direct neural transmission in a limited form—closer to binary decisions than concepts, memories or worldviews. It remains far from Egan’s Gestalt, and such interfaces may not be necessary for a useful approximation.

Semantic communication research

In telecommunications, semantic communication asks whether the meaning needed for a task survives transmission, rather than whether every symbol arrives unchanged.

Systems such as DeepSC explore neural-network-based communication measured partly by semantic similarity instead of bit-perfect reconstruction. This is not human Gestalt communication, but its question is relevant: Did the useful meaning survive?

Humans already use semantic compression

Someone familiar with the “This Is Fine” meme needs no paragraph explaining its social context. The image indexes shared knowledge; an inside joke can do the same with a word.

Technical communities have their own versions: an equation can replace pages of mathematical prose, a familiar diagram can convey system behaviour, and chess players can recognize structures in a position that require lengthy explanation to a non-player.

These are forms of semantic compression: the transmitted object is small because sender and receiver already share much of the decoder.

AI could manufacture temporary shared languages between particular people, translating between their existing conceptual structures without years spent learning the same vocabulary and shorthand.

As people worked together, their shared representations might become more compressed. A diagram that initially needed twenty minutes of explanation might eventually suffice by itself.

What might receiving a Gestalt packet feel like?

Suppose a colleague wants to explain why departments keep making inconsistent decisions despite having access to the same guidance. They provide notes, examples and a proposed remedy. Your AI turns that Gestalt packet into an explanation suited to your background. An initial view might look like this:

Central idea

Organizations often mistake making information available for creating shared understanding. Publishing the same policy does not ensure that everyone interprets it the same way.

Core causal model

  • Information availability → individual access
  • Individual access ≠ shared interpretation
  • Different interpretations → different decisions
  • Different decisions → organizational inconsistency

The sender proposes that teams compare their interpretations through examples before acting, rather than rely only on centrally published guidance.

Important assumptions

  • Participants have different backgrounds.
  • Much relevant knowledge is tacit: people know things they have not written down.
  • The organization changes faster than documentation.
  • Incentives influence interpretation.

Sender confidence

High confidence that differing interpretations contribute to inconsistent decisions; moderate confidence that the proposed discussions would resolve them.

An analogy given your software background

Distributed cache coherence: separate computers hold local copies of information that need to remain consistent. Here, teams hold different interpretations that need to be compared. The analogy has limits: people may legitimately disagree, whereas matching data copies is a technical objective.

Potential disagreement

The sender assumes central coordination has diminishing returns. Your previous arguments suggest you may favour clearer central guidance. The system would flag that possible disagreement for you to confirm or correct.

Explore

You could open a causal diagram, a worked example or counterexample, the underlying evidence, or the original source material. If the analogy did not help, you could request a different explanation.

You might then test your understanding: would publishing another policy solve the problem if the teams still interpreted it differently? Your answer could reveal what needs clarification without requiring you to agree with the proposed remedy.

Another receiver might explore the same packet through a story or organizational diagram. The transmitted object is a structured model from which an appropriate communication experience can be generated.

How would we know if it works?

Model alignment makes the idea testable. Take a complex system that Person A understands and Person B does not. Compare A explaining it through conventional writing, conventional multimedia or an AI-mediated Gestalt system. Then test whether B can:

  • explain A’s intended meaning faithfully, retaining important qualifications;
  • predict what A would predict and identify causal relationships;
  • distinguish confident claims from speculation and unresolved uncertainty;
  • identify disagreements with A and explain their basis;
  • answer counterfactual questions, recognize the model’s limits and transfer it to a new problem.

Measure time and attention through time-to-model-alignment and conceptual alignment per minute of attention. Also check factual accuracy against independent evidence, compare reconstructions with original material, and ask both participants to assess fidelity.

To test resistance to manipulation, vary analogies and persuasive cues while holding claims and evidence constant. Can B still identify the same limitations and disagreements?

Test informed consent and user control too: can participants explain what personal information is used, change or disable personalization, correct profiles and revoke consent?

Determining how to measure these outcomes reliably would itself be an HCI and cognitive-science research problem.

What about augmented reality, haptics and brain-computer interfaces?

Augmented reality could help navigate spatial representations, haptics could add perceptual dimensions, and brain-computer interfaces could eventually offer new input and output channels.

But I no longer think they are prerequisites. An initial system could run on a computer, tablet or phone. The innovation would be the semantic protocol and AI mediation layer; better interfaces could extend it later.

Risks, safeguards and design constraints

Understanding without persuasion

A decoder that knows your vocabulary, values and likely reactions could optimize for agreement instead of understanding. It should make rhetorical adaptations visible and offer alternative explanations. Understanding my model should leave you free to reject it. The system should clarify whether disagreements concern facts, assumptions, definitions, values or predictions, rather than treating consensus as the goal. Questions about agency and communication goals also feature in Hancock, Naaman and Levy’s research agenda for AI-mediated communication.

Privacy and personal control

Receiver profiles should use only necessary information, with explicit, revocable consent, clear retention limits, access controls and encryption. Users should know how their conversations and preferences are used, and be able to inspect, correct or delete their profiles. Either participant should control personalization, data use, compression and permission to infer sensitive characteristics.

Truth and faithful reconstruction

Two people can align around an inaccurate model. Gestalt packets should retain visible sources, evidence quality, competing explanations, and confidence and uncertainty labels, separating facts, interpretations, assumptions and recommendations. The sender’s confidence is not evidence: unsupported assumptions, contradictions, missing perspectives and claims needing independent verification should remain visible.

Compression can also discard something important. Replacing my biological analogy with cache coherence might lose a qualification. Receivers should be able to inspect the packet, original material, omitted details and alternative renderings, with a transformation history explaining important changes.

What still needs testing

These are design requirements, not established capabilities of a complete system. Prototypes would need tests for manipulation, omission, distortion and overconfidence, with human review for consequential uses. Detecting subtle shifts in meaning and presenting uncertainty without overwhelming the receiver remain research questions.

The governing principle should be simple: AI helps people understand one another without deciding what they must believe.

From higher bandwidth to better understanding

In 2023 I suggested lo-fi Gestalt communication might arrive soon. I would now say that pieces have arrived: multimodal AI and semantic modelling, research on human representational alignment and machine latent communication, and limited direct neural transmission.

The opportunity is to combine these ideas to help two humans align complex mental models faster and more accurately. A successful system must improve understanding while preserving truth-seeking, disagreement, privacy and autonomy.

Egan’s version remains remote. A useful approximation may not require uploading minds, direct neural interfaces or unlimited bandwidth. It may require:

an AI that understands enough about what I mean and what you already know to create an explanation that builds understanding with less unnecessary effort.

That is less spectacular than telepathy, but it may be more useful. Many components needed to begin experimenting already exist.

From concept to experiment: I am exploring this with SharedShape, an interactive thought experiment about moving the working shape of an idea between people, then testing whether it arrived intact enough to predict, decide and act.