Reference dictionary
AI & ML Glossary
Foundations, Google Cloud commands, agent development and hosting, engineering, algorithms, RAG, clustering, neural networks, and optimization.
93 of 93 terms
Core Architecture & Foundations
Mathematical and computer science foundations used to build AI systems.
- Neural Network
- A computational model loosely inspired by the brain. Layers of interconnected nodes process data and learn patterns by adjusting their parameters during training.Example: A network trained on thousands of labeled photos learns to flag whether a new X-ray shows a fracture, adjusting its internal weights each time it guesses wrong.Open Atlas lesson
- Tensor
- A multidimensional array of numbers used to represent data features and model weights. A vector is a 1D tensor; a matrix is a 2D tensor.Example: A batch of 32 color images at 224×224 pixels is one 4D tensor of shape [32, 224, 224, 3] flowing through the model.Open Atlas lesson
- Embeddings
- Numeric vectors that represent data such as words, sentences, or images. Their positions in a learned space allow algorithms to measure similarity between items.Example: The sentences “How do I reset my password?” and “I forgot my login” map to nearby vectors, so a support search returns both for either query.Open Atlas lesson
- Loss Function
- A mathematical function that measures prediction error or another training objective. Training seeks to reduce this loss; lower training loss alone does not guarantee better accuracy on new data.Example: If a model predicts a house sells for $310k and it actually sells for $300k, a squared-error loss records (10,000)² for that example and training nudges the weights to shrink it.Open Atlas lesson
- Tokenization
- Breaking text into smaller units called tokens, which may be words, characters, or subwords, so a language model can process them.Example: “unbelievable” might split into the subword tokens “un”, “believ”, and “able”, letting the model handle words it never saw whole during training.Open Atlas lesson
- Hallucination
- When a generative AI model produces content that is incorrect, unsupported, illogical, or fabricated, even when the response sounds confident.Example: Asked for sources, a chatbot cites a convincing-looking research paper with a real-sounding title and authors — but the paper does not exist.Open Atlas lesson
Google Cloud Platform (GCP) AI Services
Google services for building, testing, and deploying AI applications at scale.
- Vertex AI Custom Training Job
- A managed job that runs user-supplied training code with a configured worker pool, machine types and optional accelerators. Data access, dependencies, costs and model evaluation remain the team’s responsibility.Example: A forecasting team submits Python training code, reads job logs and compares chronological holdout error before deciding whether to register the model.Open command reference
- Vertex AI Pipelines
- Managed execution of compatible machine-learning pipeline workflows composed of steps and artifacts. A pipeline orchestrates work; it does not inherently validate data quality or guarantee an improved model.Example: A PipelineJob runs data preparation, training and holdout evaluation, with an explicit evaluation gate before deployment.Open command reference
- Vertex AI Model Registry & Endpoint
- The Model Registry manages model versions and metadata; an endpoint hosts deployed models for online prediction. Registering a model is not deployment or proof of serving-container compatibility, and deployed capacity can incur charges.Example: A team registers a candidate forecast model, tests its serving container during deployment, and routes endpoint traffic only after reviewing prediction results.Open command reference
- Document AI Processor
- A configured Document AI resource that applies a supported document-processing capability such as OCR or processor-specific entity extraction. Output fields, limits and supported modes depend on the processor and version.Example: An invoice processor returns text and invoice entities; application code follows text-anchor offsets to recover source spans and preserves currency and dimension units.Open learning example
- Cloud Vision Label Detection
- An image-analysis capability that returns descriptive labels with confidence scores. Labels describe detected content, not a custom-trained class taxonomy, and scores should not be treated as calibrated probabilities.Example: A real API request may label a bicycle photo with bicycle and wheel. The portal’s sample labels are fixed illustrations, not live Vision results.Open learning example
- Vertex AI
- Google Cloud’s managed AI platform, bringing together tools for model development, training, deployment, and machine learning operations.Example: A team trains a demand-forecasting model on Vertex AI Training, registers it in the Model Registry, and deploys it to an endpoint that serves predictions to their retail app.Open learning example
- Vertex AI Studio
- A workspace within Vertex AI for prototyping generative AI applications, testing prompts, and exploring supported models and tuning workflows.Example: A product manager drafts a customer-email summarizer by iterating on a prompt in Vertex AI Studio, comparing Gemini responses side by side before handing the prompt to engineering.Open learning example
- Model Garden
- A catalog within Vertex AI for discovering and working with Google models such as Gemini, open models such as Llama, and models from other providers. Testing and deployment options depend on the model.Example: A developer browses Model Garden, picks an open embedding model, and deploys it to a Vertex AI endpoint without provisioning any infrastructure manually.Open practice grid
- Gemini
- Google’s family of multimodal generative AI models. Supported models can work with combinations of text, code, images, audio, and video; capabilities vary by model.Example: A field technician photographs a damaged machine part and asks Gemini to identify the component and draft a repair order from the image.Open learning example
- BigQuery ML
- Machine learning capabilities in Google’s BigQuery data warehouse that let analysts create, train, and use supported models with SQL, reducing the need to move data into separate modeling tools.Example: An analyst runs CREATE MODEL in SQL to train a churn classifier directly on the customer table, then scores at-risk accounts with a SELECT query — no data export needed.Open learning example
Engineering & Deployment Frameworks
Frameworks and practices for developing and running AI applications.
- TensorFlow
- An open-source machine learning framework developed by Google for building, training, and deploying models, including deep neural networks.Example: An engineer defines a convolutional network in TensorFlow/Keras, trains it on GPUs with model.fit(), and exports a SavedModel for serving.Open Python libraries
- TensorFlow Lite (TFLite)
- A lightweight framework for running machine learning models on edge devices such as mobile phones and embedded systems. Its on-device runtime is now part of Google’s LiteRT ecosystem; microcontrollers use specialized runtimes such as TensorFlow Lite Micro.Example: A fitness app converts its pose-estimation model to TFLite so it runs in real time on the phone’s processor, even with no network connection.Open practice grid
- ONNX Runtime
- A cross-platform engine for efficient model inference, with APIs for languages including Python and C++. ONNX stands for Open Neural Network Exchange, a model format; ONNX Runtime is the execution engine.Example: A team exports their PyTorch model to ONNX and serves it from a C++ service via ONNX Runtime, cutting inference latency without rewriting the model.Open practice grid
- Retrieval-Augmented Generation (RAG)
- An architecture that retrieves relevant material from external sources, such as company documents or databases, and supplies it as context to a language model. Grounding responses this way can reduce hallucinations, but does not eliminate them.Example: An HR assistant retrieves the actual vacation policy from the company wiki and passes it to the model, so the answer quotes the real policy instead of inventing one.Open Atlas lesson
- MLOps (Machine Learning Operations)
- Engineering practices that connect model development with production operations, including repeatable training, deployment, monitoring, versioning, and retraining workflows.Example: A pipeline automatically retrains the fraud model every Sunday, runs evaluation gates, and promotes the new version to production only if precision stays above 0.95.Open command reference
Agent Development & Hosting
Agent frameworks, tool execution, evaluation and managed hosting. Local development is not a security sandbox, and hosting is distinct from enterprise publication.
- Google Agent Development Kit (ADK)
- Google’s open-source framework for building, running and evaluating agents with models, tools and orchestration. The Python package google-adk provides the adk CLI; using ADK does not automatically deploy an agent.Example: A developer creates an onboarding agent with adk create, tests draft-only tools locally, and reviews evaluation cases before choosing a hosting target.Open learning example
- Google agents-cli
- Google’s project-lifecycle tooling for scaffolding, installing, testing and deploying agent applications. It is distinct from the ADK CLI; deployment behavior depends on project configuration.Example: agents-cli create support-agent --prototype scaffolds a local prototype. A deployment-ready project can later preview its configured deployment with agents-cli deploy --dry-run.Open command reference
- Vertex AI Agent Engine
- Google Cloud’s managed services for deploying and operating agents, including a managed runtime. Some APIs and SDK resource names retain the earlier Reasoning Engine terminology. Hosted execution can incur charges and requires appropriate IAM permissions.Example: A team deploys an ADK agent to a supported region, then inspects its resource using the documented Vertex AI Python SDK instead of assuming an unverified CLI group exists.Open command reference
- Agent Tool / Function Calling
- A mechanism by which a model requests a structured call to a function or external capability. Application code validates arguments, permissions and results; a model request is not authorization to perform an action.Example: An onboarding agent requests draft_invitation with an email address. The application validates the address and creates a draft without sending it.Open learning example
- Human-in-the-loop Approval
- An explicit human decision before a sensitive action executes. Approval should cover the current action and arguments and become invalid when that proposal changes.Example: A user approves a specific invitation recipient and message. If the agent changes the recipient, the application requires new approval before sending.Open learning example
- Agent Evaluation
- Testing an agent against prepared cases, expected behavior and metrics such as response quality or tool-call correctness. Evaluations may invoke real models and tools; passing a suite is not a guarantee of safety or accuracy.Example: An ADK evaluation set checks that an unsigned onboarding agreement blocks invitation delivery, using synthetic records and a stubbed email tool.Open command reference
- Agent Session & Memory
- A session tracks one conversation’s events and state; memory can retain selected information across conversations. Persistence, retention and access controls depend on the configured service, not merely on using an agent framework.Example: A support session remembers the current ticket number. A separately configured memory service may retain an approved preference, but should not indiscriminately store private conversation text.Open Atlas lesson
- Local Agent Playground
- A development interface such as adk web or agents-cli playground for interacting with agents on a local server. Local hosting does not isolate arbitrary code or prevent paid model calls and external tool actions.Example: A developer binds the ADK web interface to 127.0.0.1, uses fictional customer records and keeps tool implementations draft-only.Open command reference
- Gemini Enterprise Agent Publication
- Making an agent available through an enterprise agent experience using the applicable integration, configuration and access policies. Hosting a runtime and publishing it to users are distinct steps; deployment alone does not guarantee enterprise visibility.Example: After hosting and testing an agent, an administrator separately configures its enterprise integration and allowed users before staff can discover it.Open command reference
Google Cloud Commands & Operations
Command-line tools, credentials, permissions and resource operations. Reference examples do not execute Cloud actions here.
- Google Cloud CLI (gcloud)
- Google’s command-line interface for managing Cloud configuration and supported resources. Commands may read data, change local settings or create billable resources depending on the operation.Example: gcloud config set project selects the active local project; gcloud services enable changes the chosen Cloud project’s API configuration.Open command reference
- BigQuery CLI (bq)
- The command-line tool for BigQuery datasets, tables, jobs and SQL queries. Query execution can incur charges; a dry run validates supported queries and estimates processing without running the query.Example: An analyst uses bq query --use_legacy_sql=false --dry_run to review a SELECT query before executing it against a large table.Open command reference
- Application Default Credentials (ADC)
- A credential-discovery mechanism used by Google client libraries. Local ADC configured with gcloud auth application-default login is separate from the credentials used by gcloud auth login; workload identity is preferable to embedded keys for hosted workloads.Example: A developer configures local ADC for a Python Document AI sample while keeping credential files out of source control.Open command reference
- Identity and Access Management (IAM)
- Google Cloud’s authorization system connecting principals, roles and resources. Authentication establishes identity; IAM determines permitted actions. API activation does not grant permissions.Example: A runtime service account can read its designated storage bucket but is not granted permission to delete unrelated datasets.Open learning example
- Service Account
- An identity used by workloads rather than an individual user. Attached identities or supported impersonation and federation avoid distributing long-lived private-key files; permissions still need least-privilege IAM roles.Example: A deployed agent uses its runtime service account to read approved documents without embedding a downloaded JSON key in its code.Open learning example
- Google Cloud Project & Region
- A project groups Cloud resources, billing association and access policies; a region identifies a geographic service location. Availability varies by service, and some APIs use locations that do not match a general region setting.Example: An Agent Engine deployment uses a supported region, while a Document AI processor uses its own returned location and matching service endpoint.Open command reference
- API Enablement
- Activating a service API for a project so permitted clients can use it. Enabling an API is a Cloud configuration change, not resource provisioning, authentication or an IAM grant.Example: gcloud services enable aiplatform.googleapis.com enables the Vertex AI API, but a deployment still needs billing, permissions and supported configuration.Open command reference
- Dry Run
- An operation-specific preview or validation mode that avoids the requested execution. Its meaning depends on the tool: a query dry run, storage-sync preview and deployment plan provide different checks and are not universal cost guarantees.Example: A team inspects agents-cli deploy --dry-run before creating infrastructure and separately reviews expected hosting charges.Open command reference
Algorithm Atlas
Learning algorithms and mechanisms explored in the interactive Algorithm Atlas.
- K-nearest neighbors (KNN)
- A method that predicts from the k closest labeled observations in feature space. Classification typically uses a majority vote; regression averages nearby target values. Feature scaling and the choice of k affect results.Example: For a new customer, KNN finds five customers with similar scaled purchase patterns. If three are repeat buyers and two are not, the majority-vote classifier predicts a repeat buyer.Open Atlas lesson
- Deep Neural Network (Spiral Dataset)
- A neural network with multiple hidden layers that compose nonlinear transformations to learn complex patterns. The Atlas uses two interlocking spiral classes to illustrate a curved decision boundary that a single straight line cannot capture.Example: A multilayer network learns to assign points to one of two spiral arms. The Atlas shows training accuracy; a separate held-out set is needed to evaluate generalization.Open Atlas lesson
- Transformer Attention
- A mechanism that uses query–key similarity scores to weight and combine value vectors, allowing tokens to incorporate context from other tokens. Softmax normalizes each query’s attention weights to sum to one.Example: In “The animal rested because it was tired”, the Atlas gives “it” a high synthetic attention weight toward “animal”. Lowering temperature concentrates the weights; these hand-designed scores do not demonstrate learned language understanding.Open Atlas lesson
- Gradient Descent & Backpropagation
- Backpropagation applies the chain rule to compute a loss’s gradients through a computational graph. Gradient descent then updates parameters in the opposite direction of those gradients, scaled by a learning rate.Example: On the Atlas loss (w₁ − 2)² + (w₂ − 3)², repeated updates with a suitable learning rate move weights toward (2, 3). An excessively large rate can overshoot or diverge.Open Atlas lesson
- Decision Trees (Gini Split)
- Models that partition observations using feature-based rules. A Gini-based classification tree chooses splits that reduce weighted class impurity; a pure node has Gini impurity zero.Example: A risk classifier splits applications at a feature threshold. The Atlas compares class mixtures on both sides and finds the feature-1 threshold with the lowest weighted Gini impurity among its candidates.Open Atlas lesson
- Random Forest
- An ensemble of decision trees trained with randomized observations and usually randomized feature subsets. Combining tree predictions can reduce variance; classification may use hard votes or averaged leaf class probabilities, depending on the implementation.Example: Seven trees evaluate a customer’s features; if five predict churn, the Atlas hard-vote ensemble predicts churn. The Python example instead averages leaf class probabilities, which need not equal the hard-vote fraction.Open Atlas lesson
- XGBoost (Gradient Boosting)
- An optimized gradient-boosted tree library that builds an additive ensemble using loss gradients, second-order information, and regularization. The Atlas illustrates the simpler squared-error boosting mechanism with residual-fitting CART trees, not the full XGBoost implementation.Example: A demand model starts with the mean target, then adds small tree corrections to its remaining errors. The Atlas tracks training MSE; the XGBoost Python example evaluates held-out RMSE to check performance on unseen observations.Open Atlas lesson
Retrieval-Augmented Generation Architectures
Core RAG structures and specialized retrieval strategies. These approaches can be combined; retrieval does not guarantee factual generation.
- Naive RAG
- A direct index → retrieve → generate pipeline supplies top-ranked source chunks to a language model. Basic retrieval can miss relevant context or include irrelevant chunks.Example: An internal FAQ retrieves a refund policy before answering a customer.Open Atlas lesson
- Advanced RAG
- Optimizes retrieval through query refinement, chunking, filtering, reranking, and context compression. Refinement and reranking can improve relevance but add latency and do not guarantee factual output.Example: Rewrite “money back” to “refund”, rerank policy chunks, then keep the relevant clauses.Open Atlas lesson
- Modular RAG
- Decouples retrieval and generation into interchangeable modules with configurable routing and composition. Modules can be composed, replaced, or looped; modularity is a design approach, not a quality guarantee.Example: Route a device error to manuals and a refund question to policy documents.Open Atlas lesson
- Hybrid RAG
- Combines lexical retrieval such as BM25 with semantic vector retrieval, then fuses or reranks results. Lexical and semantic signals complement each other. The live demo uses two lexical proxies, not embeddings or BM25.Example: Find the exact XR-42 error code while also searching for semantically similar troubleshooting instructions.Open Atlas lesson
- GraphRAG
- Uses entities and relationships, often alongside text retrieval and community summaries, to assemble connected evidence. Graph construction and coverage affect results; GraphRAG may also use vector retrieval rather than replacing it.Example: Follow Project Orion → Maya → Platform team to discover project responsibility.Open Atlas lesson
- Agentic RAG
- An agent chooses retrieval tools, inspects evidence, and may perform additional searches before answering. Tool loops need budgets, permissions, and stopping rules. The demo uses scripted decisions, not an autonomous agent.Example: Search project ownership, then query the owning team’s escalation contact.Open Atlas lesson
- Corrective RAG (CRAG)
- Grades retrieved evidence and uses correction or fallback retrieval when that evidence is weak. A retrieval grader can be wrong. The demo exposes a quality threshold; its fallback is local and never searches the web.Example: A policy question with weak local evidence triggers a secondary-source search before answering.Open Atlas lesson
- Self-RAG
- A trained model uses reflection tokens to decide when to retrieve and assess evidence relevance, support, and usefulness. Original Self-RAG needs learned reflection behavior. The demo uses a scripted support gate, not reflection tokens from a trained model.Example: A response draft is checked for supporting passages before it is accepted or revised.Open Atlas lesson
- Multi-Hop RAG
- Retrieves successive pieces of evidence, using an earlier result to form a more specific follow-up query. Each hop can propagate retrieval errors; bounded searches and source-level evidence matter.Example: Find who manages Orion, then identify that person’s team and escalation contact.Open Atlas lesson
- HyDE
- Generates a hypothetical document from a query, embeds that document, and retrieves real source documents with the resulting vector. The hypothetical text is a search aid, never evidence. The demo uses a fixed expansion and lexical matching, not generation or embeddings.Example: Draft a hypothetical refund-policy paragraph to retrieve the actual refund-policy document.Open Atlas lesson
- Simple RAG (original)
- Simple RAG is a basic retrieve-and-generate workflow, often used interchangeably with Naive RAG. Original RAG research also models probabilities over retrieved passages. Simple and Naive are overlapping labels, not universally distinct algorithms. The extractive demo does not implement original RAG training or probabilistic marginalization.Example: Retrieve refund-policy passages to ground a customer answer.Open Atlas lesson
- Simple RAG with memory
- Resolves follow-up questions using selected conversation history, then retrieves fresh source evidence. History may be stale or sensitive; it is context rather than independent evidence. The demo uses an editable local history cue, not persistent memory.Example: After discussing refunds, resolve “What receipt do I need?” using the prior refund topic.Open Atlas lesson
- Branched RAG
- Splits a question into retrieval branches and merges their evidence for synthesis. Branching is a workflow pattern, not a single standardized algorithm. Branch evidence can conflict or duplicate sources.Example: Search refund rules and shipping times separately for a combined order question.Open Atlas lesson
- Multimodal RAG
- Retrieves evidence across text, images, audio, or video with modality-aware representations. The demo searches fixed media transcripts, not pixels or audio. Real systems need appropriate encoders, indexing, and multimodal generation.Example: Combine an XR-42 manual with a photograph of its E17 display.Open Atlas lesson
- Adaptive RAG
- Routes a query to different retrieval strategies according to estimated complexity or evidence quality. The local router uses token count and conjunction heuristics, not a trained classifier; misrouting can omit needed evidence.Example: Retrieve once for refund rules but branch searches for refund and shipping.Open Atlas lesson
- Speculative RAG
- Drafts candidates from evidence subsets, then verifies or selects a candidate against its evidence. Real implementations use drafter and verifier models. The demo selects extracted passages with a lexical score; it provides no speed or correctness guarantee.Example: Compare refund candidates drawn from different source subsets and select the best-supported draft.Open Atlas lesson
Clustering Methods & Assignments
Six clustering families, their assumptions, and hard versus soft assignments.
- Centroid-based clustering · K-means
- Partitions observations around K centers. K-means uses arithmetic means; K-medoids uses actual representative observations, not means. Hard cluster IDs and centroids; clusters need interpretation, and feature scaling, initialization, and outliers affect the result.Example: Group customers using standardized spending and purchase frequency.Open Atlas lesson
- Hierarchical clustering · Agglomerative
- Builds a nested cluster tree by bottom-up merging; divisive methods instead split top-down. A cut height or desired count selects the final partition. A dendrogram and a hard partition at the selected cut. Single linkage can chain nearby points; other linkage choices produce different trees.Example: Explore product families at progressively coarser similarity levels.Open Atlas lesson
- Density-based clustering · DBSCAN
- Connects dense neighborhoods, labels reachable border points, and leaves sparse observations as noise. OPTICS explores density structure across neighborhood scales. Hard cluster IDs plus noise (−1). Noise is not proof of an anomaly; varying density and high-dimensional distances can defeat one global radius.Example: Separate curved sensor patterns while flagging isolated readings.Open Atlas lesson
- Distribution-based clustering · Gaussian mixture
- Models observations as a mixture of probability distributions; a Gaussian mixture learns component means, covariance matrices, and mixture weights. Posterior component probabilities and optional hard labels. The live EM example uses diagonal covariance and a variance floor; probabilities are model-dependent, not verified confidence.Example: Model overlapping customer populations with different spending distributions.Open Atlas lesson
- Fuzzy clustering · Fuzzy C-means
- Assigns each observation membership weights across several clusters instead of only one cluster. Fuzzy weights are not automatically calibrated probabilities. Membership weights summing to one. Larger fuzzifier m spreads weights; colors use the largest weight while the selected-point bars retain the soft assignment.Example: Represent customers whose purchase patterns overlap several market segments.Open Atlas lesson
- Graph-based clustering · Spectral
- Builds an affinity graph, embeds nodes using graph Laplacian eigenvectors, and clusters the embedding. Spectral clustering is not merely finding connected components. Hard labels from an eigenvector embedding. The live example uses a dense RBF graph and normalized Laplacian; affinity width and K affect results, and eigen-decomposition scales poorly.Example: Separate two interlocking moon-shaped groups that a straight boundary cannot separate.Open Atlas lesson
- Hard clustering
- Assigns each clustered observation to one group rather than distributing membership across groups. DBSCAN also permits unassigned noise; a hard assignment is not a known ground-truth category.Example: K-means assigns a customer to segment 2, even if the customer is almost equally close to segment 1.Open Atlas lesson
- Soft clustering
- Retains membership across multiple groups. Gaussian mixtures produce posterior probabilities under a model; fuzzy C-means produces weights that are not automatically calibrated probabilities.Example: An overlapping customer receives fuzzy memberships 0.6 and 0.4 instead of only a single segment ID.Open Atlas lesson
Neural Networks & Attention
Network architectures, QKV matching, attention relationships, and key/value sharing.
- Feedforward network · MLP
- Layers apply affine maps and nonlinear activations with no recurrent state. Without nonlinearities, stacked affine layers collapse to one affine map. Fixed 2→3→1 weights show hidden activations and a sigmoid output; the number is not a calibrated prediction or trained accuracy.Example: Predict customer churn from a fixed feature vector.Open Atlas lesson
- Convolutional network · CNN
- Shared local kernels scan spatial grids. Spatial equivariance depends on padding and stride; classification invariance is not automatic. A fixed 3×3 filter creates a 4×4 feature map from a 6×6 input; this is one cross-correlation layer, not an image classifier.Example: Detect vertical boundaries in an image.Open Atlas lesson
- Recurrent networks · RNN and LSTM
- RNNs propagate hidden state through time. LSTMs add gated cell state to help retain signals; they do not eliminate every long-range learning problem. Fixed scalar recurrence compares RNN hidden state with an LSTM-like gated cell. No sequence training or forecasting accuracy is measured.Example: Process a six-step sensor sequence.Open Atlas lesson
- Transformer · attention and feedforward
- Transformer blocks combine attention, position information, feedforward layers, residual paths, and normalization. Training is parallelizable; autoregressive generation still proceeds token by token. Fixed vectors show attention, a residual addition, and a scalar feedforward map. Position encoding, normalization, and learned projections are omitted; this is not an LLM.Example: Combine a token vector with contextual token information.Open Atlas lesson
- Autoencoder · bottleneck reconstruction
- An encoder compresses an observation and a decoder reconstructs it. A useful representation depends on the objective, capacity, and data. A fixed orthonormal linear bottleneck illustrates rank and reconstruction error, not a trained neural autoencoder. Low error does not certify normality or denoising.Example: Compress a noisy six-value sensor pattern.Open Atlas lesson
- Generative adversarial network · GAN
- A generator and discriminator optimize competing objectives. The original minimax game is zero-sum; common practical generator losses differ. One scalar generator and logistic discriminator alternate updates. This does not synthesize images, guarantee convergence, or reproduce a full GAN.Example: Shift a small synthetic distribution toward a reference distribution.Open Atlas lesson
- Query–Key–Value attention
- Queries request context, keys provide matching descriptors, and values carry content. These are learned projections in real models. Scores, normalized weights, and a weighted context vector; fixed vectors do not demonstrate language understanding.Example: Match a token query against four fixed token vectors.Open Atlas lesson
- Soft attention
- Continuously weights all unmasked source positions using a normalized distribution. Temperature changes concentration; weights are not calibrated explanations of model reasoning.Example: Blend context from several tokens.Open Atlas lesson
- Hard attention
- Selects discrete source positions rather than a continuous weighted mixture. Exact selection is not differentiable; sampling estimators, relaxations, or straight-through estimators may be used. The live argmax is deterministic inference, not REINFORCE training; changing temperature does not change the winning token.Example: Select the single strongest matching source token.Open Atlas lesson
- Self-attention
- Queries, keys, and values originate from the same sequence. A causal mask can exclude future tokens. A mask changes accessible sources; synthetic weights are not proof of learned pronoun resolution.Example: Let “it” attend to fixed vectors in its own sentence.Open Atlas lesson
- Cross-attention
- Queries come from one representation sequence and keys and values from another. The fixed target/source vectors illustrate different sequence lengths, not actual translation.Example: Match a target token to three source-language vectors.Open Atlas lesson
- Multi-head attention · MHA
- Separate query, key, and value projections provide several attention heads; outputs are concatenated and projected. Head patterns differ through fixed projections. Semantic head roles are not guaranteed; the live panel shows contexts before the learned output projection.Example: Inspect four fixed heads reading the same token sequence.Open Atlas lesson
- Grouped-query attention · GQA
- Groups of query heads share key/value heads while retaining distinct queries. Half the illustrative MHA cache scalars at equal shapes; quality and speed depend on implementation and hardware.Example: Four query heads share two key/value groups.Open Atlas lesson
- Multi-query attention · MQA
- All query heads share one key head and one value head. One quarter of illustrative MHA cache scalars at equal shapes; this is not a measured speed benchmark.Example: Four query heads read a single shared key/value representation.Open Atlas lesson
Optimization & Backpropagation
Gradient computation, data batching, momentum, adaptive scaling, and decoupled weight decay.
- Backpropagation · chain rule through a network
- Computes derivatives of the loss through a computational graph; it does not update parameters. A scalar tanh network shows exact analytical gradients; finite differences check them. Gradient computation alone leaves all weights unchanged.Example: Explain which weights contributed to a prediction error.Open Atlas lesson
- Full-dataset gradient descent
- Averages gradients over the entire training set before a parameter update. Full gradients avoid sampling noise; stability depends on curvature and learning rate. Full passes can be expensive on large datasets.Example: Fit a small linear model with all twelve observations.Open Atlas lesson
- Stochastic gradient descent · SGD
- Uses one shuffled example per update rather than a full-dataset average. One observation costs less per update, but noisy gradients can fluctuate; faster convergence and escape from minima are not guaranteed.Example: Update a sensor model one observation at a time.Open Atlas lesson
- Mini-batch gradient descent
- Averages gradients over a small subset before updating parameters. Batch size trades computation and gradient variance. The toy sample has no GPU timing or measured throughput.Example: Average three sensor observations per update in this demonstration.Open Atlas lesson
- Classical momentum
- Accumulates past gradients to influence the next parameter step. Momentum may damp some oscillations but can overshoot. It does not guarantee avoiding local minima.Example: Follow a consistent direction while fitting slope and intercept.Open Atlas lesson
- Nesterov accelerated gradient
- Computes the gradient at a momentum lookahead point before applying the next step. The fixed-rate momentum convention shown here is not every library implementation; overshoot and instability remain possible.Example: Compare a lookahead slope gradient with its current-position gradient.Open Atlas lesson
- Adagrad · accumulated squared gradients
- Divides each gradient by the square root of its accumulated squared history. Accumulation can make effective steps very small; larger steps for rare coordinates depend on their gradient history.Example: Inspect different historical scaling for slope and intercept.Open Atlas lesson
- RMSprop · decaying squared-gradient average
- Uses an exponential average of squared gradients instead of a permanently increasing sum. Forgetting old gradients avoids Adagrad’s ever-growing sum, but does not ensure nonzero steps or convergence.Example: Track how recent gradients rescale two model parameters.Open Atlas lesson
- Adam · adaptive moment estimation
- Combines exponentially averaged gradients and squared gradients with bias correction. Additional moment buffers use memory. Learning rate and other settings still need tuning; no optimizer is universally best.Example: Inspect the first and second moment estimates while fitting a line.Open Atlas lesson
- AdamW · decoupled weight decay
- Applies weight decay separately from the adaptively scaled loss gradient. Decay is not equivalent to adding L2 loss inside Adam. This toy decays both coordinates; production commonly excludes biases and normalization parameters.Example: Shrink slope and intercept separately from the Adam update.Open Atlas lesson