콘텐츠로 이동

Gemini를 사용한 컨텍스트 캐싱

ADK에서 지원Python v1.15.0Java v0.1.0Kotlin v0.7.0

에이전트를 사용하여 작업을 완료할 때 생성 AI 모델에 대한 여러 에이전트 요청에 걸쳐 확장된 지침이나 대규모 데이터 세트를 재사용하고 싶을 수 있습니다. 각 에이전트 요청에 대해 이 데이터를 다시 보내는 것은 느리고 비효율적이며 비용이 많이 들 수 있습니다. 생성 AI 모델에서 컨텍스트 캐싱 기능을 사용하면 응답 속도를 크게 높이고 각 요청에 대해 모델로 전송되는 토큰 수를 줄일 수 있습니다.

ADK 컨텍스트 캐싱 기능을 사용하면 Gemini 2.0 이상 모델을 포함하여 이를 지원하는 생성 AI 모델로 요청 데이터를 캐시할 수 있습니다. 이 문서에서는 이 기능을 구성하고 사용하는 방법을 설명합니다.

컨텍스트 캐싱 구성

에이전트를 래핑하는 ADK App 객체 수준에서 컨텍스트 캐싱 기능을 구성합니다. 다음 코드 샘플과 같이 ContextCacheConfig 클래스를 사용하여 이러한 설정을 구성합니다.

from google.adk import Agent
from google.adk.apps.app import App
from google.adk.agents.context_cache_config import ContextCacheConfig

root_agent = Agent(
  name='my_caching_agent',
  # Gemini 2.0 이상을 사용하는 에이전트 구성
)

# 컨텍스트 캐싱 구성으로 앱 만들기
app = App(
    name='my-caching-agent-app',
    root_agent=root_agent,
    context_cache_config=ContextCacheConfig(
        min_tokens=2048,    # 캐싱을 트리거하는 최소 토큰
        ttl_seconds=600,    # 최대 10분 동안 저장
        cache_intervals=5,  # 5회 사용 후 새로 고침
    ),
)
import com.google.adk.agents.BaseAgent;
import com.google.adk.agents.ContextCacheConfig;
import com.google.adk.apps.App;
import java.time.Duration;

// 컨텍스트 캐싱 구성으로 App 생성
App app = App.builder()
             .name("my-caching-agent-app")
             .rootAgent(rootAgent)
             .contextCacheConfig(
                 new ContextCacheConfig(
                     5, /* cache_intervals (최대 호출 횟수) */
                     Duration.ofMinutes(10), /* ttl */
                     2048 /* min_tokens */))
             .build();
@file:OptIn(ExperimentalContextCachingFeature::class)

import com.google.adk.kt.agents.ContextCacheConfig
import com.google.adk.kt.agents.LlmAgent
import com.google.adk.kt.annotations.ExperimentalContextCachingFeature
import com.google.adk.kt.apps.App
import com.google.adk.kt.models.Gemini
import kotlin.time.Duration.Companion.minutes

val rootAgent =
    LlmAgent(
        name = "my_caching_agent",
        // configure an agent using Gemini 2.0 or higher
        model = Gemini(name = "gemini-flash-latest"),
    )

// Create the app with context caching configuration
@OptIn(ExperimentalContextCachingFeature::class)
val app =
    App(
        appName = "my-caching-agent-app",
        rootAgent = rootAgent,
        contextCacheConfig =
            ContextCacheConfig(
                // Gemini는 모델에 따라 고유한 최소 캐시 가능 크기를 적용합니다.
                minTokens = 8192,
                ttl = 10.minutes, // 최대 10분 동안 저장
                cacheIntervals = 5, // 5회 사용 후 새로고침
                // 타임아웃 시 생성이 실패하고 요청은 캐시되지 않은 상태로 진행됩니다.
                createHttpOptions = HttpOptions(timeout = 10.seconds),
            ),
    )

구성 설정

ContextCacheConfig 클래스에는 에이전트에 대한 캐싱 작동 방식을 제어하는 다음 설정이 있습니다. 이러한 설정을 구성하면 앱 내의 모든 에이전트에 적용됩니다.

  • min_tokens (int): 캐싱을 활성화하는 데 필요한 요청의 최소 토큰 수입니다. 이 설정을 사용하면 성능 이점이 미미한 매우 작은 요청에 대한 캐싱 오버헤드를 피할 수 있습니다. 기본값은 0입니다.
  • ttl_seconds (int): 캐시의 TTL(Time-To-Live) (초)입니다. 이 설정은 캐시된 콘텐츠가 새로 고쳐지기 전에 저장되는 기간을 결정합니다. 기본값은 1800(30분)입니다.
  • cache_intervals (int): 동일한 캐시된 콘텐츠가 만료되기 전에 사용할 수 있는 최대 횟수입니다. 이 설정을 사용하면 TTL이 만료되지 않았더라도 캐시가 업데이트되는 빈도를 제어할 수 있습니다. 기본값은 10입니다.
  • create_http_options (HttpOptions): 캐시 생성 호출에 대한 HTTP 옵션으로, 타임아웃을 설정할 수 있습니다. 호출 시간이 초과되면 실패하고 요청은 캐싱 없이 진행됩니다. Python 및 Kotlin에서 사용 가능하며 기본값은 없음(none)입니다.

캐시 사용 여부 확인

ADK에서 지원Kotlin v0.6.0

캐싱이 활성화되면 LLM 응답을 기반으로 하는 이벤트는 해당 호출에 대해 캐시가 수행한 작업을 보고하는 CacheMetadata를 포함할 수 있습니다. 캐싱이 비활성화되었거나 호출에서 캐시 정보를 생성하지 않은 경우 null이므로 읽기 전에 확인하세요. 존재하는 경우 두 가지 상태가 있습니다. cacheName, expireTime, invocationsUsed가 모두 설정된 활성 캐시(active cache) 상태와 세 가지가 모두 null인 지문 전용(fingerprint-only) 상태입니다.

/** Reports whether the context cache was used for the LLM call behind [event]. */
fun logCacheUse(event: Event) {
    // Null when caching is disabled, and on any event whose LLM call produced
    // no cache information.
    val cache = event.cacheMetadata ?: return

    if (!cache.isActive) {
        // Fingerprint-only: ADK measured the cacheable prefix but no cache is in
        // use. That is the first turn, a prefix that changed since the last turn,
        // or a cache ADK did not create -- most often because the cacheable
        // prefix was below minTokens.
        println("Not cached yet; fingerprinted ${cache.contentsCount} contents.")
        return
    }

    println("Cache ${cache.cacheName} reused ${cache.invocationsUsed} time(s).")
    if (cache.expireSoon) {
        // Advisory only. ADK goes on reusing the cache until it actually expires,
        // so this is a heads-up for your own code, not a prediction about the
        // next turn.
        println("Cache is at or near expiry.")
    }
}

expireSoon은 캐시가 약 2분 이내에 만료되거나 이미 만료되었음을 의미합니다. 이는 사용자 코드에 대한 신호일 뿐 ADK가 직접 조치를 취하는 것은 아닙니다. ADK는 실제로 expireTime이 지나거나 cacheIntervals를 초과하거나 캐시된 접두사가 변경될 때까지 캐시를 계속 재사용합니다.

토큰 수는 CacheMetadata에 포함되지 않으므로 LlmResponse.usageMetadata에서 읽으세요.

다음 단계

컨텍스트 캐싱 기능을 사용하고 테스트하는 방법에 대한 전체 구현은 다음 샘플을 참고하세요.

  • cache_analysis: 컨텍스트 캐싱의 성능을 분석하는 방법을 보여주는 코드 샘플입니다.

사용 사례에서 세션 전체에서 사용되는 지침을 제공해야 하는 경우 에이전트에 대한 static_instruction 매개변수를 사용하는 것을 고려하세요. 이 매개변수를 사용하면 생성 모델에 대한 시스템 지침을 수정할 수 있습니다. 자세한 내용은 다음 샘플 코드를 참고하세요.

  • static_instruction: 정적 지침을 사용하는 디지털 펫 에이전트의 구현입니다.