📚 AI Tutorials

30 tutorials. Each solves a specific problem with steps, prompt examples, and tips. From beginner to pro.

1
Guide

How to Use AI Tools to Create Viral Xiaohongshu Graphic Posts: A Complete Workflow

2
Guide

Cloud Run Instances for Long-Lived Agents: When a $5.70/Month Runtime Beats a VM

Google Cloud introduced Cloud Run instances in preview on August 27, 2026. The product targets a gap between autoscaling Cloud Run services and fully managed virtual machines: workloads that want exactly one long-lived container, stable state, a persistent HTTPS endpoint, and occasional CPU bursts. A Cloud Run instance has no autoscaling, runs one instance, can run continuously for up to seven days before restart, retains a stable HTTPS URL across updates and restarts, and can be stopped and resumed. Google’s published example prices a continuously running 1-vCPU, 1-GiB instance at about $5.70 for 30 days. The model fits personal agents and low-duty-cycle workers, but it is not a replacement for stateless autoscaling services, Kubernetes, or full VMs.

3
Guide

Qwen3 Embedding on Cloud TPU: Production Long-Context Retrieval with vLLM

Google Cloud published native vLLM TPU support for embedding inference on August 26, 2026, targeting production retrieval rather than chat generation. The engineering work focuses on Qwen3-Embedding-8B and Qwen3-VL-Embedding-8B with long text and multimodal contexts, including 16K-class text sequences and 15K+ multimodal inputs. Google addressed TPU tensor alignment, lazy loading, JAX/XLA compilation warm-up, chunked prefill, and pooling-state preservation through a hybrid StepPool design. In one published Qwen3-Embedding-8B configuration using bf16, 16K+ sequences, and TP=4, TPU Ironwood reached 83,996 total tokens/s and 5.13 requests/s. Google also validates cross-hardware vector parity with cosine-similarity thresholds of at least 0.999 for text and 0.995 for multimodal inputs.

4
Guide

Copilot Code Review for Azure Repos: How to Configure It for Real Engineering Value

Microsoft announced the public preview of GitHub Copilot Code Review for Azure Repos on August 26, 2026. Azure DevOps customers no longer need early-access registration. The preview includes enterprise controls that matter in practice: organization/project/repository enablement, Managed DevOps Pools, custom instructions at multiple scopes, automatic review through branch policies, draft pull-request review, and project-level cost attribution in Azure Cost Management. The value is not simply generating more comments. The value comes from inserting a consistent automated first review into a governed engineering pipeline.

5
Guide

WebMCP in Practice: Turning Websites into Agent-Callable Applications

WebMCP is an experimental open web standard that lets a website expose structured tools directly to AI agents. Instead of relying on screenshots, DOM guessing, selectors, and simulated clicks, an agent can discover a semantic operation such as `get_order_status` with a clear description, JSON input schema, and execution function. Chrome provides both imperative and declarative approaches, with an origin trial beginning in Chrome 149. OpenAI opened a 10-day WebMCP Challenge on August 25, and ChatGPT’s in-app browser can test WebMCP-enabled sites. The bigger product shift is that websites may soon need to design not only human UX, but agent UX.

6
Guide

Codex mcp-server Is Deprecated: A Practical App Server Migration Guide

On August 24, 2026, OpenAI deprecated the `codex mcp-server` command. The replacement depends on the actual integration goal. Deep product integrations should use the Codex App Server; CI, batch jobs, and non-interactive automation should generally use the Codex SDK; and users invoking Codex from Claude Code should use the Codex plugin for Claude Code. App Server is not simply MCP Server under a new name. It is a bidirectional Codex control protocol exposing authentication, threads, turns, approvals, streamed agent events, MCP integration, configuration, and runtime state through JSON-RPC-style messages.

7
Guide

ADK Live Voice Eval: Turning “Sounds Good” into a Testable CI Gate

Google introduced native live evaluation for Agent Development Kit on August 24. Developers can now drive a live voice agent with an LLM-controlled simulated user that speaks through synthesized audio, score the resulting multi-turn trajectory with natural-language rubrics and per-turn metrics, and run the same evaluation loop programmatically in CI/CD. Google’s example uses three `gemini-live-2.5-flash-native-audio` agents in a graph workflow, `gemini-3.7-flash` for simulated-user turn logic, and `gemini-3.1-flash-tts-preview` for audio generation. This moves voice-agent testing beyond manual demos into repeatable regression engineering.

8
Guide

Running a Local Agent on Raspberry Pi 5 with LiteRT and Gemma

Google recently demonstrated a practical Edge AI stack for Raspberry Pi 5 using LiteRT and lightweight Gemma models. For Gemma 4 E2B, Google reports roughly 99 tokens/s prefill, 9 tokens/s decode, and about 1,432 MB peak memory on Raspberry Pi 5. In the Reachy Mini voice demo, effective generation reached about 27.3 characters per second, approximately 300 English words per minute. More importantly, LiteRT can orchestrate more than an LLM: CPU resources can handle Gemma and speech recognition while the Pi GPU continuously runs vision models such as YOLO. This makes it possible to build private, offline multimodal agents without a cloud dependency.

9
Guide

Zero-Trust Agents with Google ADK: Where Security Boundaries Should Live

On August 17, 2026, Google published a zero-trust reference architecture for agents built with Agent Development Kit (ADK). The motivation is straightforward: once an agent can issue refunds, modify databases, execute dynamic code, and call internal APIs, “do not perform dangerous actions” inside a system prompt is no longer a real security boundary. Google’s proposed pattern emphasizes three hard controls: cryptographic identity for state-changing writes, sandboxed execution for generated code, and deterministic semantic gateways that validate business meaning before actions reach production systems. This article translates those ideas into an enterprise architecture.

10
Guide

ChatGPT Sites in Practice: From Prompt to Production Site

ChatGPT Sites is currently in public beta, and OpenAI’s documentation has just been updated with clearer guidance on creation, preview, publishing, sharing, custom domains, and workspace permissions. Sites is more than generating HTML. Users can describe an interactive website or lightweight application inside ChatGPT Work or desktop Codex, review a private preview, save versions, publish a production URL, and control access. Current use cases include dashboards, project trackers, launch calendars, internal portals, calculators, prototypes, and lightweight reports. This guide focuses on a safe production workflow rather than a feature list.

11
Guide

Google Credentio + C2PA: Local Verification for AI Images, Video, Audio, and Documents

On August 13, 2026, Google open-sourced Credentio, a high-performance C++ library for validating C2PA Content Credentials, initially supporting specification versions 2.2 and 2.4. Google says the underlying code already powers nearly 40 conformant C2PA-enabled Google products and has scaled to tens of billions of generated assets across images, video, audio, documents, and other file formats. The most important design feature is local-first validation: media does not need to be uploaded to Google or another remote validation service, reducing privacy, bandwidth, latency, and file-size problems. This article explains how C2PA differs from AI detectors and model watermarks, and provides practical architectures for CMS upload pipelines, digital asset management, desktop applications, and enterprise trust policies.

12
Guide

Google Cloud API Gateway Model Routing: One OpenAI-Compatible Endpoint for Multiple LLMs

Enterprise AI applications increasingly use more than one model. Low-cost models may handle routine support, stronger models may handle reasoning or coding, and open models may serve batch workloads. If every application integrates directly with every provider, credentials, retries, rate limits, model IDs, and cost tracking spread across the codebase. Google Cloud API Gateway now offers Model Routing in Public Preview. It can expose an OpenAI-compatible endpoint and route requests, through OpenAPI 3.x configuration, to Google-hosted Gemini, Claude, or OpenAI OSS-GPT backends. The real value is not calling three models through one API. It is decoupling business code from model vendors.

13
Guide

GitHub Copilot for JetBrains Enterprise Governance: MCP Allowlists, Plugins, and OpenTelemetry

On August 18, 2026, GitHub added enterprise-managed settings to GitHub Copilot for JetBrains. Administrators can now centrally govern plugin enablement, approved plugin marketplaces, MCP server access, OpenTelemetry collection, and permission modes such as Bypass Approvals or Autopilot. This is more significant than a normal IDE update. Once coding agents can edit files, execute commands, invoke MCP tools, and interact with enterprise systems, local developer settings become part of the organization’s security boundary.

14
Guide

Gemini Live API Async Function Calling for Real-Time Voice Agents

One of the biggest usability problems in real-time voice agents is tool latency. If the model needs to query CRM, orders, databases, search, or ticketing systems and every function call blocks the conversation, the user experiences several seconds of silence. Gemini Live API supports asynchronous function calling in its cascaded architecture. A function can be marked `NON_BLOCKING`, allowing the live conversation to continue while the tool executes in the background. This does not mean the model can invent tool-dependent facts. It means developers need explicit task state, timeouts, cancellation, concurrency controls, result correlation, permissions, and idempotency.

15
Guide

OpenAI’s “Defender’s Window”: What Enterprises Should Do Now

On August 17, 2026, OpenAI published “The Defender’s Window,” arguing that frontier AI is beginning to automate meaningful parts of real-world cyber operations. Long-standing weaknesses—software flaws, forgotten permissions, misconfigurations, leaked credentials, and security debt—are becoming cheaper to discover and combine. The same capabilities can help defenders review code, triage alerts, enumerate attack paths, prioritize vulnerability backlogs, generate patches, and conduct investigations. The “defender’s window” is the period in which organizations can use capable AI to improve their defenses before similar capabilities become more broadly available to attackers.

16
Guide

OpenAI Assistants API Shuts Down August 26: Complete Responses API Migration Guide

OpenAI will shut down the Assistants API on August 26, 2026. Applications still using Assistants, Threads, Runs, Run Steps, File Search, or Assistant-based function calling are now inside the final migration window. The Responses API has reached feature parity and is the primary direction for OpenAI agent development, with newer capabilities such as Conversations, MCP, computer use, and deep research. This guide covers concept mapping, state migration, tool loops, prompt versioning, file search, shadow traffic, rollout, observability, and rollback.

17
Guide

AI Assistant Memory Architecture

AI memory is not permanent storage of every conversation. Short-term context, preferences, task state, episodic records, and semantic knowledge require different storage, retrieval, update, and deletion policies.

18
Guide

Building a Personal AI Knowledge System

The most common personal knowledge problem is growing collections that are never reused. This guide covers capture, curation, sources, structure, summaries, links, retrieval, review, and deletion.

19
Guide

OCR and Document AI Architecture

Enterprise documents include native PDFs, scans, images, tables, handwriting, and Office files. A single OCR path cannot handle them all. This guide defines classification, routing, multiple parsers, quality checks, and a unified representation.

20
Guide

AI Invoice and Document Extraction Pipeline

Invoice, order, and receipt automation requires more than OCR text. It needs layout, schemas, validation, duplicate detection, confidence, human review, and auditing.

21
Guide

AI Music Copyright Risk Guide

AI music rights involve platform terms, user inputs, training disputes, lyrics, voice likeness, similarity, regional law, and distribution-platform rules.

22
Guide

Commercial AI Music Production Workflow

Commercial AI music production cannot stop at generation and download. Teams need a complete process for briefs, style references, lyrics, versions, mixing, review, terms, and rights evidence.

23
Guide

Voice Cloning Compliance Guide

Voice is a highly identifying personal characteristic. Enterprise cloning requires consent scope, identity verification, use restrictions, disclosure, storage, third-party access, and withdrawal.

24
Guide

AI Podcast Production Workflow

AI can accelerate podcast research, scripting, voice, cleanup, chapters, and distribution, but it cannot guarantee facts, pacing, or rights.

25
Guide

Multitenant AI SaaS Architecture

Multitenancy in AI SaaS extends beyond database rows to vector indexes, object storage, caches, prompts, tool credentials, model quotas, and logs.

26
Guide

Building an AI SaaS MVP Backend

Even an AI SaaS MVP must handle users, tenants, model credentials, streaming, usage, quotas, and retries correctly. This guide defines a minimal but extensible backend.

27
Guide

Deploying Local LLMs on Office PCs

Running local models on office PCs requires matching model size and quantization to RAM, VRAM, CPU, operating system, and tasks rather than chasing parameter count.

28
Guide

AI Portal Authorization: SSO, RBAC, Multitenancy, and Data Isolation

An AI portal connects models, knowledge, and tools. Authorization failures can leak data across departments or enable unauthorized actions. This guide defines layered identity, role, attribute, resource, tenant, retrieval, and tool controls.

29
Guide

Deploying an Internal Enterprise AI Portal

An enterprise AI portal needs a unified entrance, but the chat UI should not connect directly to model providers. This guide covers identity, gateways, knowledge permissions, quotas, auditing, safety, rollout, and adoption.

30
Guide

Source Credibility Scoring for AI Research

AI search often returns many citations, but citation count does not equal evidence quality. This guide scores source authority, independence, freshness, applicability, and claim support.