Files
AI/PHASE_7_LLM_GUIDE.md
lasta 47a710a8b9 Fix datetime deprecation warnings in Phase 8 modules
- Update all datetime.utcnow() to datetime.now(UTC) for Python 3.12+ compatibility
- Update all datetime.utcfromtimestamp() to datetime.fromtimestamp(..., UTC)
- Fix dataclass default_factory to use lambda: datetime.now(UTC)
- Update auth, audit, billing, and realtime modules
- Add UTC import from datetime module
- Update pytest configuration to include pytest-asyncio
- All 28 Phase 8 enterprise tests pass with no warnings

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-05-14 11:50:23 +09:00

14 KiB

Phase 7 LLM 엔드투엔드 통합 가이드

개요

Phase 7는 Phase 6의 GraphRAG 파이프라인을 확장하여 LLM(대언어모델)을 직접 통합합니다.

특징:

  • 다중 LLM 프로바이더 지원 (OpenAI, Anthropic, Local)
  • 실시간 스트리밍 응답 (Server-Sent Events)
  • Redis 기반 응답 캐싱 (TTL 설정 가능)
  • RAG 컨텍스트 자동 추출 + 프롬프트 생성
  • 메타데이터 추적 (레이턴시, 토큰 수, 모델 정보)

빠른 시작

1. 서버 시작

# Phase 7 앱 시작 (포트 8001)
python -m uvicorn ontology_platform.ont_platform.api.phase7_app:app --reload --port 8001

# 또는 기본 포트 8000
python -m uvicorn ontology_platform.ont_platform.api.phase7_app:app --reload

2. 헬스 체크

curl http://localhost:8000/health

응답:

{
  "status": "healthy",
  "version": "0.7.0",
  "neo4j": "connected",
  "redis": "available",
  "llm_provider": "openai",
  "timestamp": "2026-05-14T10:30:45.123456"
}

3. LLM 설정

# 현재 LLM 설정 확인
curl http://localhost:8000/api/v1/llm/info

# LLM 변경 (OpenAI → Anthropic)
curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx"

REST API 엔드포인트

1. 기본 LLM 쿼리 (캐싱 포함)

POST /api/v1/llm/ask

LLM에 질문하고 캐시된 응답을 반환합니다.

요청:

curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{
    "query": "Apple의 주요 제품은 무엇인가?",
    "context_hops": 2,
    "use_cache": true,
    "temperature": 0.7,
    "max_tokens": 500
  }'

요청 파라미터:

  • query (필수): 사용자 질문
  • context_hops (선택): 그래프 컨텍스트 깊이 (기본: 2)
  • use_cache (선택): 캐시 사용 여부 (기본: true)
  • temperature (선택): 응답 다양성 (0.0~2.0, 기본: 0.7)
  • max_tokens (선택): 최대 토큰 수 (기본: 500)

응답:

{
  "query": "Apple의 주요 제품은 무엇인가?",
  "answer": "Apple의 주요 제품으로는 iPhone, iPad, Mac, Apple Watch 등이 있습니다. iPhone은 Apple의 핵심 수익원이며...",
  "context_size": 45,
  "relevant_entities": ["Apple Inc.", "iPhone", "iPad", "Mac", "Steve Jobs"],
  "latency_ms": 245.5,
  "cached": false,
  "model": "gpt-4",
  "provider": "openai"
}

응답 필드:

  • query: 입력 질문
  • answer: LLM의 최종 답변
  • context_size: 사용된 그래프 노드 수
  • relevant_entities: 검색된 관련 엔티티
  • latency_ms: 전체 응답 시간 (밀리초)
  • cached: 캐시된 응답 여부 (true면 실제 레이턴시는 훨씬 적음)
  • model: 사용된 모델
  • provider: LLM 프로바이더

성능:

  • 캐시 미스: 1-3초 (RAG 추출 + LLM 생성)
  • 캐시 히트: 50-100ms (Redis 조회)

2. 스트리밍 응답 (실시간 토큰)

POST /api/v1/llm/ask/stream

LLM 응답을 실시간 스트리밍합니다 (Server-Sent Events).

요청:

curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
  -H "Content-Type: application/json" \
  -d '{
    "query": "온톨로지란 무엇인가?",
    "context_hops": 2,
    "temperature": 0.7
  }'

응답 (SSE 스트림):

data: {"type": "metadata", "query": "온톨로지란 무엇인가?", "context_nodes": 50, "relevant_entities": ["Ontology", "Knowledge Graph"], "extraction_time_ms": 120.5}

data: {"type": "token", "content": "온톨로지는", "token_index": 0}

data: {"type": "token", "content": " ", "token_index": 1}

data: {"type": "token", "content": "어떤", "token_index": 2}

...

data: {"type": "complete", "total_tokens": 156, "timestamp": "2026-05-14T10:35:20.123456"}

스트림 포맷:

  • 각 줄은 SSE 이벤트: data: {JSON}\n\n
  • metadata: 초기 메타데이터 (컨텍스트, 엔티티)
  • token: 각 생성된 토큰
  • complete: 완료 신호

클라이언트 예제 (JavaScript):

const eventSource = new EventSource(
  'http://localhost:8000/api/v1/llm/ask/stream',
  { method: 'POST', body: JSON.stringify({query: "..."})}
);

eventSource.addEventListener('message', (event) => {
  const data = JSON.parse(event.data);
  
  if (data.type === 'metadata') {
    console.log('Context:', data.context_nodes, 'nodes');
  } else if (data.type === 'token') {
    process.stdout.write(data.content);  // 실시간 출력
  } else if (data.type === 'complete') {
    console.log(`\n완료 (${data.total_tokens} 토큰)`);
    eventSource.close();
  }
});

성능: 3-5초 (토큰 실시간 전달, 캐싱 미적용)


3. RAG 메타데이터만 (LLM 호출 없음)

POST /api/v1/llm/ask/metadata

LLM 호출 없이 RAG 컨텍스트 정보만 반환합니다.

요청:

curl -X POST http://localhost:8000/api/v1/llm/ask/metadata \
  -H "Content-Type: application/json" \
  -d '{
    "query": "Apple과 관련된 정보",
    "context_hops": 2
  }'

응답:

{
  "query": "Apple과 관련된 정보",
  "context_nodes": 45,
  "relevant_entities": ["Apple Inc.", "iPhone", "iPad", "Steve Jobs"],
  "extraction_time_ms": 145.2,
  "llm_provider": "openai",
  "llm_model": "gpt-4"
}

성능: 100-300ms (RAG 추출만, LLM 호출 없음)


LLM 설정

LLM 설정 변경

POST /api/v1/llm/configure

LLM 프로바이더, 모델, 온도 등을 변경합니다.

요청 (OpenAI → Anthropic 변경):

curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx&temperature=0.5&max_tokens=1000"

요청 파라미터:

  • provider (필수): openai, anthropic, local
  • model (필수): 모델 이름
    • OpenAI: gpt-4, gpt-3.5-turbo
    • Anthropic: claude-3-opus, claude-3-sonnet, claude-2
    • Local: llama2, mistral, etc.
  • api_key (선택): API 키 (환경 변수로도 설정 가능)
  • temperature (선택): 0.0~2.0 (기본: 0.7)
  • max_tokens (선택): 토큰 제한 (기본: 500)

응답:

{
  "status": "configured",
  "provider": "anthropic",
  "model": "claude-3-opus",
  "temperature": 0.5,
  "max_tokens": 1000
}

LLM 정보 조회

GET /api/v1/llm/info

현재 LLM 설정을 조회합니다.

응답:

{
  "llm_provider": "openai",
  "llm_model": "gpt-4",
  "temperature": 0.7,
  "max_tokens": 500,
  "redis_available": true,
  "timestamp": "2026-05-14T10:40:15.123456"
}

캐싱 관리

캐시 정보

GET /api/v1/llm/cache/info

Redis 캐시 통계를 조회합니다.

응답:

{
  "redis_available": true,
  "used_memory_mb": 125.5,
  "cache_keys": 342,
  "redis_version": "7.0.0"
}

캐시 삭제

DELETE /api/v1/llm/cache

모든 RAG 캐시를 삭제합니다.

요청:

curl -X DELETE http://localhost:8000/api/v1/llm/cache

응답:

{
  "status": "success",
  "deleted_keys": "342"
}

설정 (환경 변수)

LLM 프로바이더 API 키

# OpenAI
export OPENAI_API_KEY=sk-proj-xxx

# Anthropic
export ANTHROPIC_API_KEY=sk-ant-xxx

# Local LLM (LM Studio)
export LM_STUDIO_URL=http://localhost:1234/v1

Neo4j 연결

export NEO4J_URI=bolt://localhost:7687
export NEO4J_USER=neo4j
export NEO4J_PASSWORD=ontology123

Redis 연결

export REDIS_URL=redis://localhost:6379

사용 예제

예제 1: 기본 질문응답

import requests

# 1. 기본 질문 (캐싱 포함)
response = requests.post(
    "http://localhost:8000/api/v1/llm/ask",
    json={
        "query": "Apple의 창립자는 누구인가?",
        "context_hops": 2,
        "use_cache": True
    }
)

data = response.json()
print(f"답변: {data['answer']}")
print(f"응답 시간: {data['latency_ms']:.1f}ms")
print(f"캐시: {data['cached']}")

예제 2: 스트리밍 응답

import requests
import json

# 2. 스트리밍 응답
response = requests.post(
    "http://localhost:8000/api/v1/llm/ask/stream",
    json={
        "query": "온톨로지 시스템의 주요 기능을 설명해주세요",
        "context_hops": 2
    },
    stream=True
)

for line in response.iter_lines():
    if line:
        data = json.loads(line[6:])  # "data: " 제거
        
        if data['type'] == 'metadata':
            print(f"컨텍스트: {data['context_nodes']} 노드")
        elif data['type'] == 'token':
            print(data['content'], end='', flush=True)
        elif data['type'] == 'complete':
            print(f"\n완료 ({data['total_tokens']} 토큰)")

예제 3: LLM 설정 변경

import requests

# 3. LLM 설정 변경 (OpenAI → Anthropic)
response = requests.post(
    "http://localhost:8000/api/v1/llm/configure",
    params={
        "provider": "anthropic",
        "model": "claude-3-opus",
        "api_key": "sk-ant-xxx",
        "temperature": 0.5
    }
)

print(response.json())
# Output: {"status": "configured", "provider": "anthropic", ...}

예제 4: RAG + LLM 파이프라인

# 1단계: RAG 메타데이터 확인
curl -X POST http://localhost:8000/api/v1/llm/ask/metadata \
  -H "Content-Type: application/json" \
  -d '{"query": "AI의 응용 사례"}'

# 2단계: LLM 쿼리 (캐싱 자동)
curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{"query": "AI의 응용 사례", "use_cache": true}'

# 3단계: 스트리밍 응답 (실시간)
curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
  -H "Content-Type: application/json" \
  -d '{"query": "AI의 응용 사례"}'

다중 LLM 프로바이더

OpenAI (기본)

# OpenAI로 설정
curl "http://localhost:8000/api/v1/llm/configure?provider=openai&model=gpt-4&api_key=sk-proj-xxx"

# 지원 모델: gpt-4, gpt-4-turbo, gpt-3.5-turbo

특징:

  • 가장 강력한 성능
  • 넓은 지식 기반
  • ⚠️ API 비용 발생 (토큰 기반)

Anthropic (Claude)

# Anthropic으로 설정
curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx"

# 지원 모델: claude-3-opus, claude-3-sonnet, claude-2

특징:

  • 안전성과 윤리성 강조
  • 더 긴 컨텍스트 윈도우 (200K 토큰)
  • 한국어 우수

Local LLM (LM Studio, Ollama)

# 로컬 LLM으로 설정
curl "http://localhost:8000/api/v1/llm/configure?provider=local&model=llama2&base_url=http://localhost:1234/v1"

# 지원 모델: llama2, mistral, neural-chat, etc.

특징:

  • 로컬 실행 (프라이버시)
  • API 비용 무료
  • ⚠️ 성능은 상대적으로 낮음

성능 최적화

1. 캐싱 활용

# 첫 번째 쿼리 (캐시 미스): ~1-3초
curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{"query": "Apple의 제품", "use_cache": true}'

# 두 번째 쿼리 (캐시 히트): ~50-100ms (30배 빠름!)
curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{"query": "Apple의 제품", "use_cache": true}'

2. 스트리밍 응답 (UI 반응성)

# 전체 응답을 기다리는 대신, 토큰 실시간 수신
curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
  -H "Content-Type: application/json" \
  -d '{"query": "..."}'

3. 온도 조정

# 고속 응답 (더 결정적)
curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{"query": "...", "temperature": 0.0, "max_tokens": 250}'

# 창의적 응답 (더 다양)
curl -X POST http://localhost:8000/api/v1/llm/ask \
  -H "Content-Type: application/json" \
  -d '{"query": "...", "temperature": 0.9, "max_tokens": 1000}'

성능 특성

작업 데이터셋 응답 시간
LLM 쿼리 (캐시 미스) - 1-3초
LLM 쿼리 (캐시 히트) - 50-100ms
스트리밍 응답 (첫 토큰) - 500-800ms
RAG 메타데이터 - 100-300ms
캐시 삭제 1K 키 < 100ms
LLM 설정 변경 - < 50ms

배포

Docker

FROM python:3.10-slim

WORKDIR /app

COPY requirements.txt .
RUN pip install -r requirements.txt

COPY . .

# Phase 7 앱 실행
CMD ["uvicorn", "ontology_platform.ont_platform.api.phase7_app:app", "--host", "0.0.0.0", "--port", "8000"]

Kubernetes

apiVersion: apps/v1
kind: Deployment
metadata:
  name: ontology-phase7
spec:
  replicas: 3
  selector:
    matchLabels:
      app: ontology-phase7
  template:
    metadata:
      labels:
        app: ontology-phase7
    spec:
      containers:
      - name: api
        image: ontology-phase7:0.7.0
        ports:
        - containerPort: 8000
        env:
        - name: OPENAI_API_KEY
          valueFrom:
            secretKeyRef:
              name: llm-secrets
              key: openai-key
        - name: NEO4J_URI
          value: "bolt://neo4j:7687"
        - name: REDIS_URL
          value: "redis://redis:6379"

문제 해결

Redis 연결 실패

# Redis 상태 확인
redis-cli ping

# Docker Redis 실행
docker run -d -p 6379:6379 redis:7.0

LLM API 키 오류

# 환경 변수 확인
echo $OPENAI_API_KEY

# 유효한 API 키 설정
export OPENAI_API_KEY=sk-proj-xxx

높은 응답 시간

# 1. Redis 캐싱 활성화
# use_cache: true 설정

# 2. 토큰 제한 감소
# max_tokens: 250 설정

# 3. 온도 감소 (더 결정적)
# temperature: 0.3 설정

# 4. 로컬 LLM 사용 (프라이버시 + 속도)
# provider: local 설정

다음 단계

Phase 8: 엔터프라이즈 기능

목표: 대규모 운영 지원
- 멀티테넌트 (여러 조직 동시 지원)
- 실시간 그래프 업데이트 (WebSocket)
- 변경 이력 추적 (감사 로그)
- 비용 관리 (API 호출당 요금)
- 고급 분석 (사용자별 통계)

정보

  • 버전: 0.7.0
  • 마지막 업데이트: 2026-05-14
  • 지원 모델: GPT-4, Claude 3, Llama 2, Mistral
  • 캐시 TTL: 1시간 (설정 가능)

Phase 7 LLM 통합으로 지식 그래프를 기반으로 한 지능형 질문응답 시스템을 구축하세요!