Playbooks
Tutorials
Long-form guides optimized for engineers shipping GenAI features responsibly.
12 min read
Offline vs Online Eval Frequency
This tutorial discusses the trade-offs between offline and online evaluation frequencies for machine learning models, focusing on their impact on model performance and user experience.
18 min read
Planner–Executor Loops and Failure Recovery
This tutorial explains the planner-executor loop in AI systems and how to implement effective failure recovery strategies. Prerequisites include knowledge of AI planning algorithms and system design.
20 min read
Agent Memory: Scratchpad vs Vector Store
This tutorial compares scratchpad memory and vector store memory in AI agents, focusing on their use cases and performance characteristics. Prerequisites include a basic understanding of AI memory architectures.
15 min read
Runbooks When Quality Regresses Overnight
This tutorial outlines how to create effective runbooks to address overnight quality regressions in software systems. Prerequisites include familiarity with incident management and basic scripting skills.
16 min read
Canary Prompts for Regression Detection
Utilizing canary prompts to detect regressions in language models. Prerequisites include familiarity with regression testing and LLM evaluation metrics.
22 min read
Prompt Injection Defenses in Multi-Tenant Apps
Developing strategies to protect multi-tenant applications from prompt injection attacks. Prerequisites include understanding of security vulnerabilities and multi-tenant architecture.
18 min read
Observability: Traces for LLM + Tool Spans
Implementing observability practices to trace interactions between large language models (LLMs) and external tools. Prerequisites include knowledge of observability tools and LLM architectures.
20 min read
Sandboxing Tools with Least Privilege
Implementing sandboxing techniques to limit tool access and enhance security. Prerequisites include familiarity with security protocols and system architecture.
15 min read
Human-in-the-Loop for High-Stakes Actions
Integrating human oversight in automated systems to ensure accuracy and accountability in critical scenarios. Prerequisites include understanding of automation frameworks and risk management principles.
20 min read
Graph RAG for Entity-Heavy Domains
Explore the use of Graph Retrieval-Augmented Generation (RAG) for domains with complex entities, requiring knowledge of graph databases and RAG techniques.
10 min read
Hybrid Search: BM25 + Dense Re-Ranking
This tutorial explores the integration of BM25 and dense re-ranking techniques to enhance search accuracy. Prerequisites include familiarity with information retrieval concepts and basic machine learning.
6 min read
PII Handling in Retrieval Pipelines
Effective handling of Personally Identifiable Information (PII) is essential in retrieval systems to ensure compliance and user trust. Prerequisites include knowledge of data privacy regulations and retrieval system architecture.
6 min read
Cost Controls: Batching vs Streaming Tokens
Understanding the trade-offs between batching and streaming token processing can optimize costs in NLP applications. Prerequisites include familiarity with tokenization and processing pipelines.
18 min read
Golden-Set Design for RAG Faithfulness
Understand how to design a golden set for evaluating the faithfulness of Retrieval-Augmented Generation (RAG) models. Prerequisites include familiarity with RAG systems and evaluation metrics.
15 min read
Embedding Drift Monitoring in Production
Learn how to implement embedding drift monitoring in production systems to ensure model reliability. Prerequisites include familiarity with machine learning models and data pipelines.
11 min read
Shadow Traffic for Safe Model Rollouts
Learn how to implement shadow traffic techniques for safely rolling out new models without impacting user experience.
14 min read
Reducing Hallucinations with Citation Constraints
This tutorial discusses strategies to minimize hallucinations in AI outputs by implementing citation constraints.
12 min read
Structured Outputs vs JSON Mode Tradeoffs
Explore the trade-offs between using structured outputs and JSON mode in APIs, focusing on performance and usability.
15 min read
Evaluating Tool-Calling Reliability Under Load
This tutorial provides a framework for assessing the reliability of tool calls in high-load scenarios, ensuring system robustness.
10 min read
Latency Budgets for Streaming Chat UX
Learn how to define and manage latency budgets to enhance user experience in streaming chat applications, ensuring responsiveness and engagement.
25 min read
Cross-Encoder Re-Rankers at Scale
Understand how to implement cross-encoder re-rankers for large-scale information retrieval systems. Prerequisites include knowledge of ranking algorithms and machine learning.
20 min read
Quantization Impact on Retrieval Quality
Explore the effects of quantization on the quality of information retrieval systems. Prerequisites include familiarity with machine learning models and retrieval systems.
15 min read
Synthetic Data for Classifier Fine-Tunes
Learn how to generate synthetic data to improve classifier performance, especially in scenarios with limited labeled data. Prerequisites include basic understanding of machine learning and data generation techniques.
16 min read
Chunking Strategies for Legal/Medical PDFs
Learn effective chunking strategies for processing legal and medical PDFs to enhance information retrieval. Prerequisites include familiarity with PDF processing and natural language processing concepts.