콘텐츠로 이동

런타임 구성(Runtime Configuration)

ADK에서 지원Python v0.1.0TypeScript v0.2.0Go v0.1.0Java v0.1.0Kotlin v0.1.0

RunConfig는 streaming mode, speech setting, LLM call limit, live agent option 등 에이전트의 런타임 동작을 제어합니다. 기본 동작을 재정의하려면 RunConfig를 runner.run_async() 또는 runner.run_live()에 전달합니다.

from google.adk.agents.run_config import RunConfig, StreamingMode

config = RunConfig(
    streaming_mode=StreamingMode.SSE,
    max_llm_calls=200,
)

async for event in runner.run_async(
    ...,
    run_config=config,
):
    ...
import { RunConfig, StreamingMode } from '@google/adk';

const config: RunConfig = {
  streamingMode: StreamingMode.SSE,
  maxLlmCalls: 200,
};
import "google.golang.org/adk/v2/agent"

config := agent.RunConfig{
    StreamingMode: agent.StreamingModeSSE,
}
import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;

RunConfig config = RunConfig.builder()
    .streamingMode(StreamingMode.SSE)
    .maxLlmCalls(200)
    .build();
val config =
    RunConfig(
        streamingMode = StreamingMode.SSE,
        // Cap the LLM calls a single run may make. Defaults to 500.
        maxLlmCalls = 200,
    )

// Pass it to runner.runAsync
// runner.runAsync(..., runConfig = config)

세션 및 컨텍스트 관리

ADK에서 지원Python

Long-running session에서는 얼마나 많은 history를 load할지, context window를 압축할지 제어할 수 있습니다.

  • get_session_config: session을 load할 때 가져오는 event를 제한합니다. 매 invocation마다 전체 event history를 load하지 않도록 num_recent_events 또는 after_timestamp를 사용합니다. 이러한 필터는 저장된 이벤트를 삭제하지 않고 로드되는 뷰만 제한합니다. 새 이벤트는 저장된 히스토리에 추가되며, 로드된 뷰에서 제외된 이전 이벤트는 그대로 보존됩니다.
  • context_window_compression: LLM input에 대한 context window compression을 활성화합니다. session이 model context limit에 가까워질 때 유용합니다.
  • model_input_context: 이번 호출에 대해서만 LLM 요청에 추가되는 types.Content 목록입니다. 러너는 이를 세션에 유지하지 않으므로 대화 기록을 변경하지 않고 턴별 컨텍스트를 제공할 수 있습니다.
  • include_thoughts_from_other_agents: 다른 에이전트의 생각(thought) 파트가 LLM 컨텍스트에 포함되는지 여부를 제어합니다. 기본적으로 비활성화되어 있습니다.
from google.adk.agents.run_config import RunConfig
from google.adk.sessions.base_session_service import GetSessionConfig

config = RunConfig(
    get_session_config=GetSessionConfig(num_recent_events=50),
)

텍스트 응답 옵션

에이전트가 텍스트 모드에서 응답하는 방식을 단어가 생성되는 대로 전달할지, 아니면 하나의 전체 응답으로 전달할지 아래에 설명된 스트리밍 모드(Streaming Mode) 매개변수로 제어할 수 있습니다:

  • StreamingMode.NONE(기본값): runner가 turn마다 하나의 완성된 응답을 반환합니다. CLI tool, batch processing, synchronous workflow에 적합합니다.
  • StreamingMode.SSE: Server-Sent Events streaming입니다. LLM이 생성하는 동안 runner가 partial event를 yield하여 typewriter-style UI와 real-time chat display를 구현할 수 있습니다.

음성 입력과 출력을 포함한 데이터의 양방향 스트리밍을 활성화하는 스트리밍 모드 매개변수의 또 다른 설정이 있습니다. 이 기능은 단순한 에이전트 이상의 추가 구성이 필요합니다. 이 기능에 대한 자세한 내용은 라이브 및 음성 에이전트를 참고하세요.

StreamingMode.SSE와 함께 support_cfc=True를 설정하면 Compositional Function Calling(CFC)을 활성화할 수 있습니다. CFC는 모델이 function call을 동적으로 구성하고 실행할 수 있게 하며, 내부적으로 Live API를 사용합니다.

Experimental

CFC 지원은 실험적이며 향후 릴리스에서 API나 동작이 변경될 수 있습니다.

from google.adk.agents.run_config import RunConfig, StreamingMode

config = RunConfig(
    streaming_mode=StreamingMode.SSE,
    support_cfc=True,
    max_llm_calls=150,
)
import { RunConfig, StreamingMode } from '@google/adk';

const config: RunConfig = {
    streamingMode: StreamingMode.SSE,
    maxLlmCalls: 150,
};
import "google.golang.org/adk/v2/agent"

config := agent.RunConfig{
    StreamingMode: agent.StreamingModeSSE,
}
import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;

RunConfig config = RunConfig.builder()
    .streamingMode(StreamingMode.SSE)
    .maxLlmCalls(150)
    .build();
// Note: Kotlin currently has no supportCfc equivalent
val streamingConfig =
    RunConfig(
        streamingMode = StreamingMode.SSE,
        maxLlmCalls = 150,
    )

오디오 및 음성 구성

ADK에서 지원PythonTypeScriptJava

Voice-enabled agent에서는 speech synthesis, audio transcription, response modality를 구성합니다.

라이브 에이전트

이 섹션에서는 언어 전반에 공통으로 사용되는 오디오 필드를 다룹니다. 전사 스트리밍, 음성 선택, 음성 활동 감지(VAD), 능동적/감정적 대화 등 전체 라이브(run_live()) 구성 참조는 라이브 에이전트 구성을 참고하세요.

  • speech_config: 음성 출력의 음성과 언어를 설정합니다(예: en-US와 "Kore" 음성).
  • response_modalities: 출력 형식을 제어합니다. 세션은 정확히 하나의 모달리티만 허용하므로, 음성 에이전트에는 ["AUDIO"]를, 텍스트 전용 에이전트에는 ["TEXT"]를 사용합니다. 음성과 텍스트를 모두 얻으려면 ["AUDIO"]로 설정하고 출력 오디오 전사(transcription)에서 텍스트를 읽으세요.
  • output_audio_transcription / input_audio_transcription: 모델의 오디오 출력 및 사용자의 오디오 입력 전사를 활성화합니다. Python에서는 둘 다 기본값이 AudioTranscriptionConfig()입니다.
from google.adk.agents.run_config import RunConfig, StreamingMode
from google.genai import types

config = RunConfig(
    speech_config=types.SpeechConfig(
        language_code="en-US",
        voice_config=types.VoiceConfig(
            prebuilt_voice_config=types.PrebuiltVoiceConfig(
                voice_name="Kore"
            )
        ),
    ),
    response_modalities=["AUDIO"],
    streaming_mode=StreamingMode.SSE,
    max_llm_calls=1000,
)
import { RunConfig, StreamingMode } from '@google/adk';
import { Modality } from '@google/genai';

const config: RunConfig = {
    speechConfig: {
        languageCode: "en-US",
        voiceConfig: {
            prebuiltVoiceConfig: {
                voiceName: "Kore"
            }
        },
    },
    responseModalities: [Modality.AUDIO],
    streamingMode: StreamingMode.SSE,
    maxLlmCalls: 1000,
};
import com.google.adk.agents.RunConfig;
import com.google.adk.agents.RunConfig.StreamingMode;
import com.google.common.collect.ImmutableList;
import com.google.genai.types.Modality;
import com.google.genai.types.PrebuiltVoiceConfig;
import com.google.genai.types.SpeechConfig;
import com.google.genai.types.VoiceConfig;

RunConfig runConfig =
    RunConfig.builder()
        .streamingMode(StreamingMode.SSE)
        .maxLlmCalls(1000)
        .responseModalities(ImmutableList.of(new Modality(Modality.Known.AUDIO)))
        .speechConfig(
            SpeechConfig.builder()
                .voiceConfig(
                    VoiceConfig.builder()
                        .prebuiltVoiceConfig(
                            PrebuiltVoiceConfig.builder().voiceName("Kore").build())
                        .build())
                .languageCode("en-US")
                .build())
        .build();

라이브 에이전트 구성

ADK에서 지원PythonTypeScriptJava

ADK 에이전트는 대화형 에이전트 경험을 생성하기 위해 라이브 및 음성 에이전트를 지원할 수 있습니다. runner.run_live() 메서드를 사용하여 이 기능을 지원하는 에이전트를 구성합니다. 라이브 에이전트(run_live()) 세션은 realtime_input_config, session_resumption, save_live_blob, tool_thread_pool_config, proactivity, enable_affective_dialog 등을 포함한 실시간 매개변수 집합을 추가합니다. 자세한 내용은 라이브 에이전트 문서를 참조하세요:

tool_thread_pool_config 설정은 예외입니다. 이는 Live API 관심사라기보다는 런타임 관심사이므로 여기에 유지됩니다. 이벤트 루프가 사용자의 인터럽트에 계속 응답할 수 있도록 도구 실행을 백그라운드 스레드 풀에서 실행합니다. 모든 매개변수가 모든 언어에서 제공되는 것은 아닙니다. 언어별 세부 정보는 API reference를 참고하세요.

from google.adk.agents.run_config import RunConfig, ToolThreadPoolConfig

config = RunConfig(
    save_live_blob=True,
    tool_thread_pool_config=ToolThreadPoolConfig(max_workers=8),
)

Thread pool and the GIL

Thread pool은 blocking I/O와 GIL을 release하는 C extension(예: time.sleep(), network call, numpy)에 도움이 됩니다. 순수 Python CPU-bound code에는 도움이 되지 않습니다. GIL이 Python bytecode의 진정한 병렬 실행을 막기 때문입니다.

import { RunConfig } from '@google/adk';

const config: RunConfig = {
    enableAffectiveDialog: true,
    proactivity: {
        proactiveAudio: true,
    },
};
import com.google.adk.agents.RunConfig;
import com.google.genai.types.AvatarConfig;

RunConfig config = RunConfig.builder()
    .avatarConfig(
        AvatarConfig.builder()
            .avatarName("PREBUILT_AVATAR_ID")
            .build())
    .build();

런타임 제한 및 디버깅 구성

다음 매개변수로 runtime guardrail과 debugging을 제어합니다.

  • max_llm_calls: run당 총 LLM call 수를 제한합니다(기본값: 500). 0 또는 음수는 무제한 호출을 의미하지만 production에서는 권장하지 않습니다. 사용하는 언어의 가장 큰 정수(Python에서는 sys.maxsize, Kotlin에서는 Int.MAX_VALUE)를 전달하면 오류가 발생합니다.
  • save_input_blobs_as_artifacts: True이면 input blob(예: uploaded file)을 debugging과 auditing용 run artifact로 저장합니다.
  • custom_metadata: invocation에 첨부되는 임의 metadata의 dict[str, Any]입니다. tracing 또는 logging에 유용합니다.

API reference

전체 field, type, default 목록은 각 언어의 API reference를 참고하세요.