cyberivy
Google ResearchGboardFederated LearningPrivacyTrusted Execution EnvironmentDifferential PrivacySigstoreOpen Source AI

Gboard trains word predictions with verifiable privacy

October 4, 2026

Nahaufnahme eines Infineon-Sicherheitschips auf einer grünen Computerplatine

Google is moving federated learning into isolated server environments. Public logs and reproducible source code are meant to make it auditable which programs access encrypted device data.

What this is about

Google Research introduced a new federated-learning architecture on October 4, 2026 that is already used for English and Japanese next-word prediction in Gboard. Its goal is not merely to keep raw data away from devices. User devices and outside auditors should also be able to determine which programs are allowed to process encrypted training data.

This matters because federated learning retained a trust problem: data could be collected in a decentralized way or combined securely, yet outsiders could not fully verify what happened on the servers. Google's approach combines Trusted Execution Environments, or TEEs, with access policies, a public transparency log, and reproducible builds.

What the new system actually does

A participating device encrypts training examples locally and specifies before upload which computing programs may access them. These rules are published in Rekor, a public transparency log from the Sigstore project. A key-management system releases decryption keys only to isolated server workloads whose identity and code match the approved policy.

Processing runs inside TEEs. These isolated environments are intended to shield the content and state of a computation from the rest of the server while providing technical attestation of which code is running. Google says operators can see only metrics and model weights protected with differential privacy. Published source code for core components is intended to support reproducible builds.

Unlike older versions, part of the gradient computation moves from the phone to servers. Google can therefore collect uploads and run the computation later in parallel. The company says this shortens training runs that previously could take one to two months, but it does not state a generally applicable new duration.

Why it matters

Federated learning already supports everyday features such as word prediction, suggested replies, and text selection. Devices provide especially sensitive signals for these functions: what people type, when they use a feature, and which corrections they make can reveal a great deal about personal habits. Technical access control is therefore stronger than a provider's promise alone.

The approach shifts the audit question from “Do you trust Google?” to “Do the published policy, attested code, and reproducible build match?” It does not eliminate trust, but it creates concrete checkpoints for security researchers and auditors. Its production use in Gboard also shows that this is more than a laboratory concept.

In plain language

Think of the system as a sealed kitchen. Ingredients arrive in locked containers. The kitchen opens them only when the recipe, cook, and workflow appear on a publicly visible list. What leaves the kitchen is not an individual dish, but a statistically protected summary of many preparations.

The seal does not prove that the recipe is sensible. Its main purpose is to show that the announced recipe was executed and that nobody secretly looked inside the containers.

A practical example

Suppose 100,000 devices provide encrypted examples of which word follows “Good.” Each device authorizes only a published training program. The key-management system checks the isolated environment's attestation before releasing data.

The program computes an updated model, applies the statistical protection required for differential privacy, and releases only protected weights and metrics. An auditor can then check whether the approved policy appeared in the transparency log and whether the executed binary can be reproduced from the published source code. The auditor cannot see individual inputs, and Google's blog post does not reveal the specific privacy budget for this hypothetical example.

Scope and limits

First, TEEs are not invulnerable vaults. Google itself points to known limitations such as side channels; flaws in hardware, firmware, attestation, or key management can undermine the security assumptions.

Second, “verifiable” does not automatically mean “independently verified.” Public code and logs create the opportunity for scrutiny. Whether enough qualified third parties continuously perform that work remains open.

Third, the architecture protects the intended data path, but it does not judge the purpose, fairness, or quality of the trained model. A correctly attested computation can still rely on a problematic policy. Google also provides neither a complete independent security assessment nor universal performance figures for other products.

SEO & GEO keywords

Google Research, Gboard, federated learning, Trusted Execution Environment, TEE, differential privacy, Rekor, Sigstore, Confidential Federated Compute, privacy, word prediction, auditable machine learning

💡 In plain English

Google trains Gboard word predictions partly on encrypted device data inside isolated server environments. Public access rules and reproducible code are meant to make the processing software auditable. This improves accountability but does not remove hardware risks or the need for independent review.

Key Takeaways

  • →Gboard already uses the architecture for English and Japanese next-word prediction.
  • →Devices encrypt examples and pre-authorize permitted server workloads.
  • →Access policies appear in the public Rekor transparency log.
  • →Only attested TEE workloads receive keys for processing.
  • →Side channels, hardware flaws, and limited outside auditing remain risks.

FAQ

Are keyboard inputs sent to Google in plain text?

Google says participating devices encrypt training examples before upload. Only attested workloads in isolated environments may process them, but the post is not an independent audit of every Gboard data transfer.

What is a Trusted Execution Environment?

A TEE is an isolated area of a computer designed to protect code and data from the rest of the system and attest which software it runs.

Is the system completely secure?

No. Side channels and flaws in hardware, firmware, attestation, or key management remain possible attack paths.

Can outsiders inspect the source code?

Google publishes core Confidential Federated Compute components. Reproducible builds are intended to let auditors compare source code with the running workload.

Sources & Context