- Update all datetime.utcnow() to datetime.now(UTC) for Python 3.12+ compatibility - Update all datetime.utcfromtimestamp() to datetime.fromtimestamp(..., UTC) - Fix dataclass default_factory to use lambda: datetime.now(UTC) - Update auth, audit, billing, and realtime modules - Add UTC import from datetime module - Update pytest configuration to include pytest-asyncio - All 28 Phase 8 enterprise tests pass with no warnings Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
646 lines
14 KiB
Markdown
646 lines
14 KiB
Markdown
# Phase 7 LLM 엔드투엔드 통합 가이드
|
|
|
|
## 개요
|
|
|
|
Phase 7는 Phase 6의 GraphRAG 파이프라인을 확장하여 **LLM(대언어모델)을 직접 통합**합니다.
|
|
|
|
**특징**:
|
|
- ✅ 다중 LLM 프로바이더 지원 (OpenAI, Anthropic, Local)
|
|
- ✅ 실시간 스트리밍 응답 (Server-Sent Events)
|
|
- ✅ Redis 기반 응답 캐싱 (TTL 설정 가능)
|
|
- ✅ RAG 컨텍스트 자동 추출 + 프롬프트 생성
|
|
- ✅ 메타데이터 추적 (레이턴시, 토큰 수, 모델 정보)
|
|
|
|
---
|
|
|
|
## 빠른 시작
|
|
|
|
### 1. 서버 시작
|
|
|
|
```bash
|
|
# Phase 7 앱 시작 (포트 8001)
|
|
python -m uvicorn ontology_platform.ont_platform.api.phase7_app:app --reload --port 8001
|
|
|
|
# 또는 기본 포트 8000
|
|
python -m uvicorn ontology_platform.ont_platform.api.phase7_app:app --reload
|
|
```
|
|
|
|
### 2. 헬스 체크
|
|
|
|
```bash
|
|
curl http://localhost:8000/health
|
|
```
|
|
|
|
응답:
|
|
```json
|
|
{
|
|
"status": "healthy",
|
|
"version": "0.7.0",
|
|
"neo4j": "connected",
|
|
"redis": "available",
|
|
"llm_provider": "openai",
|
|
"timestamp": "2026-05-14T10:30:45.123456"
|
|
}
|
|
```
|
|
|
|
### 3. LLM 설정
|
|
|
|
```bash
|
|
# 현재 LLM 설정 확인
|
|
curl http://localhost:8000/api/v1/llm/info
|
|
|
|
# LLM 변경 (OpenAI → Anthropic)
|
|
curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx"
|
|
```
|
|
|
|
---
|
|
|
|
## REST API 엔드포인트
|
|
|
|
### 1. 기본 LLM 쿼리 (캐싱 포함)
|
|
|
|
#### `POST /api/v1/llm/ask`
|
|
|
|
LLM에 질문하고 **캐시된 응답**을 반환합니다.
|
|
|
|
**요청**:
|
|
```bash
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"query": "Apple의 주요 제품은 무엇인가?",
|
|
"context_hops": 2,
|
|
"use_cache": true,
|
|
"temperature": 0.7,
|
|
"max_tokens": 500
|
|
}'
|
|
```
|
|
|
|
**요청 파라미터**:
|
|
- `query` (필수): 사용자 질문
|
|
- `context_hops` (선택): 그래프 컨텍스트 깊이 (기본: 2)
|
|
- `use_cache` (선택): 캐시 사용 여부 (기본: true)
|
|
- `temperature` (선택): 응답 다양성 (0.0~2.0, 기본: 0.7)
|
|
- `max_tokens` (선택): 최대 토큰 수 (기본: 500)
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"query": "Apple의 주요 제품은 무엇인가?",
|
|
"answer": "Apple의 주요 제품으로는 iPhone, iPad, Mac, Apple Watch 등이 있습니다. iPhone은 Apple의 핵심 수익원이며...",
|
|
"context_size": 45,
|
|
"relevant_entities": ["Apple Inc.", "iPhone", "iPad", "Mac", "Steve Jobs"],
|
|
"latency_ms": 245.5,
|
|
"cached": false,
|
|
"model": "gpt-4",
|
|
"provider": "openai"
|
|
}
|
|
```
|
|
|
|
**응답 필드**:
|
|
- `query`: 입력 질문
|
|
- `answer`: LLM의 최종 답변
|
|
- `context_size`: 사용된 그래프 노드 수
|
|
- `relevant_entities`: 검색된 관련 엔티티
|
|
- `latency_ms`: 전체 응답 시간 (밀리초)
|
|
- `cached`: 캐시된 응답 여부 (true면 실제 레이턴시는 훨씬 적음)
|
|
- `model`: 사용된 모델
|
|
- `provider`: LLM 프로바이더
|
|
|
|
**성능**:
|
|
- 캐시 미스: 1-3초 (RAG 추출 + LLM 생성)
|
|
- 캐시 히트: 50-100ms (Redis 조회)
|
|
|
|
---
|
|
|
|
### 2. 스트리밍 응답 (실시간 토큰)
|
|
|
|
#### `POST /api/v1/llm/ask/stream`
|
|
|
|
LLM 응답을 **실시간 스트리밍**합니다 (Server-Sent Events).
|
|
|
|
**요청**:
|
|
```bash
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"query": "온톨로지란 무엇인가?",
|
|
"context_hops": 2,
|
|
"temperature": 0.7
|
|
}'
|
|
```
|
|
|
|
**응답 (SSE 스트림)**:
|
|
```
|
|
data: {"type": "metadata", "query": "온톨로지란 무엇인가?", "context_nodes": 50, "relevant_entities": ["Ontology", "Knowledge Graph"], "extraction_time_ms": 120.5}
|
|
|
|
data: {"type": "token", "content": "온톨로지는", "token_index": 0}
|
|
|
|
data: {"type": "token", "content": " ", "token_index": 1}
|
|
|
|
data: {"type": "token", "content": "어떤", "token_index": 2}
|
|
|
|
...
|
|
|
|
data: {"type": "complete", "total_tokens": 156, "timestamp": "2026-05-14T10:35:20.123456"}
|
|
```
|
|
|
|
**스트림 포맷**:
|
|
- 각 줄은 SSE 이벤트: `data: {JSON}\n\n`
|
|
- `metadata`: 초기 메타데이터 (컨텍스트, 엔티티)
|
|
- `token`: 각 생성된 토큰
|
|
- `complete`: 완료 신호
|
|
|
|
**클라이언트 예제 (JavaScript)**:
|
|
```javascript
|
|
const eventSource = new EventSource(
|
|
'http://localhost:8000/api/v1/llm/ask/stream',
|
|
{ method: 'POST', body: JSON.stringify({query: "..."})}
|
|
);
|
|
|
|
eventSource.addEventListener('message', (event) => {
|
|
const data = JSON.parse(event.data);
|
|
|
|
if (data.type === 'metadata') {
|
|
console.log('Context:', data.context_nodes, 'nodes');
|
|
} else if (data.type === 'token') {
|
|
process.stdout.write(data.content); // 실시간 출력
|
|
} else if (data.type === 'complete') {
|
|
console.log(`\n완료 (${data.total_tokens} 토큰)`);
|
|
eventSource.close();
|
|
}
|
|
});
|
|
```
|
|
|
|
**성능**: 3-5초 (토큰 실시간 전달, 캐싱 미적용)
|
|
|
|
---
|
|
|
|
### 3. RAG 메타데이터만 (LLM 호출 없음)
|
|
|
|
#### `POST /api/v1/llm/ask/metadata`
|
|
|
|
LLM 호출 **없이** RAG 컨텍스트 정보만 반환합니다.
|
|
|
|
**요청**:
|
|
```bash
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask/metadata \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"query": "Apple과 관련된 정보",
|
|
"context_hops": 2
|
|
}'
|
|
```
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"query": "Apple과 관련된 정보",
|
|
"context_nodes": 45,
|
|
"relevant_entities": ["Apple Inc.", "iPhone", "iPad", "Steve Jobs"],
|
|
"extraction_time_ms": 145.2,
|
|
"llm_provider": "openai",
|
|
"llm_model": "gpt-4"
|
|
}
|
|
```
|
|
|
|
**성능**: 100-300ms (RAG 추출만, LLM 호출 없음)
|
|
|
|
---
|
|
|
|
## LLM 설정
|
|
|
|
### LLM 설정 변경
|
|
|
|
#### `POST /api/v1/llm/configure`
|
|
|
|
LLM 프로바이더, 모델, 온도 등을 변경합니다.
|
|
|
|
**요청 (OpenAI → Anthropic 변경)**:
|
|
```bash
|
|
curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx&temperature=0.5&max_tokens=1000"
|
|
```
|
|
|
|
**요청 파라미터**:
|
|
- `provider` (필수): `openai`, `anthropic`, `local`
|
|
- `model` (필수): 모델 이름
|
|
- OpenAI: `gpt-4`, `gpt-3.5-turbo`
|
|
- Anthropic: `claude-3-opus`, `claude-3-sonnet`, `claude-2`
|
|
- Local: `llama2`, `mistral`, etc.
|
|
- `api_key` (선택): API 키 (환경 변수로도 설정 가능)
|
|
- `temperature` (선택): 0.0~2.0 (기본: 0.7)
|
|
- `max_tokens` (선택): 토큰 제한 (기본: 500)
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"status": "configured",
|
|
"provider": "anthropic",
|
|
"model": "claude-3-opus",
|
|
"temperature": 0.5,
|
|
"max_tokens": 1000
|
|
}
|
|
```
|
|
|
|
### LLM 정보 조회
|
|
|
|
#### `GET /api/v1/llm/info`
|
|
|
|
현재 LLM 설정을 조회합니다.
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"llm_provider": "openai",
|
|
"llm_model": "gpt-4",
|
|
"temperature": 0.7,
|
|
"max_tokens": 500,
|
|
"redis_available": true,
|
|
"timestamp": "2026-05-14T10:40:15.123456"
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## 캐싱 관리
|
|
|
|
### 캐시 정보
|
|
|
|
#### `GET /api/v1/llm/cache/info`
|
|
|
|
Redis 캐시 통계를 조회합니다.
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"redis_available": true,
|
|
"used_memory_mb": 125.5,
|
|
"cache_keys": 342,
|
|
"redis_version": "7.0.0"
|
|
}
|
|
```
|
|
|
|
### 캐시 삭제
|
|
|
|
#### `DELETE /api/v1/llm/cache`
|
|
|
|
모든 RAG 캐시를 삭제합니다.
|
|
|
|
**요청**:
|
|
```bash
|
|
curl -X DELETE http://localhost:8000/api/v1/llm/cache
|
|
```
|
|
|
|
**응답**:
|
|
```json
|
|
{
|
|
"status": "success",
|
|
"deleted_keys": "342"
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## 설정 (환경 변수)
|
|
|
|
### LLM 프로바이더 API 키
|
|
|
|
```bash
|
|
# OpenAI
|
|
export OPENAI_API_KEY=sk-proj-xxx
|
|
|
|
# Anthropic
|
|
export ANTHROPIC_API_KEY=sk-ant-xxx
|
|
|
|
# Local LLM (LM Studio)
|
|
export LM_STUDIO_URL=http://localhost:1234/v1
|
|
```
|
|
|
|
### Neo4j 연결
|
|
|
|
```bash
|
|
export NEO4J_URI=bolt://localhost:7687
|
|
export NEO4J_USER=neo4j
|
|
export NEO4J_PASSWORD=ontology123
|
|
```
|
|
|
|
### Redis 연결
|
|
|
|
```bash
|
|
export REDIS_URL=redis://localhost:6379
|
|
```
|
|
|
|
---
|
|
|
|
## 사용 예제
|
|
|
|
### 예제 1: 기본 질문응답
|
|
|
|
```python
|
|
import requests
|
|
|
|
# 1. 기본 질문 (캐싱 포함)
|
|
response = requests.post(
|
|
"http://localhost:8000/api/v1/llm/ask",
|
|
json={
|
|
"query": "Apple의 창립자는 누구인가?",
|
|
"context_hops": 2,
|
|
"use_cache": True
|
|
}
|
|
)
|
|
|
|
data = response.json()
|
|
print(f"답변: {data['answer']}")
|
|
print(f"응답 시간: {data['latency_ms']:.1f}ms")
|
|
print(f"캐시: {data['cached']}")
|
|
```
|
|
|
|
### 예제 2: 스트리밍 응답
|
|
|
|
```python
|
|
import requests
|
|
import json
|
|
|
|
# 2. 스트리밍 응답
|
|
response = requests.post(
|
|
"http://localhost:8000/api/v1/llm/ask/stream",
|
|
json={
|
|
"query": "온톨로지 시스템의 주요 기능을 설명해주세요",
|
|
"context_hops": 2
|
|
},
|
|
stream=True
|
|
)
|
|
|
|
for line in response.iter_lines():
|
|
if line:
|
|
data = json.loads(line[6:]) # "data: " 제거
|
|
|
|
if data['type'] == 'metadata':
|
|
print(f"컨텍스트: {data['context_nodes']} 노드")
|
|
elif data['type'] == 'token':
|
|
print(data['content'], end='', flush=True)
|
|
elif data['type'] == 'complete':
|
|
print(f"\n완료 ({data['total_tokens']} 토큰)")
|
|
```
|
|
|
|
### 예제 3: LLM 설정 변경
|
|
|
|
```python
|
|
import requests
|
|
|
|
# 3. LLM 설정 변경 (OpenAI → Anthropic)
|
|
response = requests.post(
|
|
"http://localhost:8000/api/v1/llm/configure",
|
|
params={
|
|
"provider": "anthropic",
|
|
"model": "claude-3-opus",
|
|
"api_key": "sk-ant-xxx",
|
|
"temperature": 0.5
|
|
}
|
|
)
|
|
|
|
print(response.json())
|
|
# Output: {"status": "configured", "provider": "anthropic", ...}
|
|
```
|
|
|
|
### 예제 4: RAG + LLM 파이프라인
|
|
|
|
```bash
|
|
# 1단계: RAG 메타데이터 확인
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask/metadata \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "AI의 응용 사례"}'
|
|
|
|
# 2단계: LLM 쿼리 (캐싱 자동)
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "AI의 응용 사례", "use_cache": true}'
|
|
|
|
# 3단계: 스트리밍 응답 (실시간)
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "AI의 응용 사례"}'
|
|
```
|
|
|
|
---
|
|
|
|
## 다중 LLM 프로바이더
|
|
|
|
### OpenAI (기본)
|
|
|
|
```bash
|
|
# OpenAI로 설정
|
|
curl "http://localhost:8000/api/v1/llm/configure?provider=openai&model=gpt-4&api_key=sk-proj-xxx"
|
|
|
|
# 지원 모델: gpt-4, gpt-4-turbo, gpt-3.5-turbo
|
|
```
|
|
|
|
**특징**:
|
|
- ✅ 가장 강력한 성능
|
|
- ✅ 넓은 지식 기반
|
|
- ⚠️ API 비용 발생 (토큰 기반)
|
|
|
|
### Anthropic (Claude)
|
|
|
|
```bash
|
|
# Anthropic으로 설정
|
|
curl "http://localhost:8000/api/v1/llm/configure?provider=anthropic&model=claude-3-opus&api_key=sk-ant-xxx"
|
|
|
|
# 지원 모델: claude-3-opus, claude-3-sonnet, claude-2
|
|
```
|
|
|
|
**특징**:
|
|
- ✅ 안전성과 윤리성 강조
|
|
- ✅ 더 긴 컨텍스트 윈도우 (200K 토큰)
|
|
- ✅ 한국어 우수
|
|
|
|
### Local LLM (LM Studio, Ollama)
|
|
|
|
```bash
|
|
# 로컬 LLM으로 설정
|
|
curl "http://localhost:8000/api/v1/llm/configure?provider=local&model=llama2&base_url=http://localhost:1234/v1"
|
|
|
|
# 지원 모델: llama2, mistral, neural-chat, etc.
|
|
```
|
|
|
|
**특징**:
|
|
- ✅ 로컬 실행 (프라이버시)
|
|
- ✅ API 비용 무료
|
|
- ⚠️ 성능은 상대적으로 낮음
|
|
|
|
---
|
|
|
|
## 성능 최적화
|
|
|
|
### 1. 캐싱 활용
|
|
|
|
```bash
|
|
# 첫 번째 쿼리 (캐시 미스): ~1-3초
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "Apple의 제품", "use_cache": true}'
|
|
|
|
# 두 번째 쿼리 (캐시 히트): ~50-100ms (30배 빠름!)
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "Apple의 제품", "use_cache": true}'
|
|
```
|
|
|
|
### 2. 스트리밍 응답 (UI 반응성)
|
|
|
|
```bash
|
|
# 전체 응답을 기다리는 대신, 토큰 실시간 수신
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask/stream \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "..."}'
|
|
```
|
|
|
|
### 3. 온도 조정
|
|
|
|
```bash
|
|
# 고속 응답 (더 결정적)
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "...", "temperature": 0.0, "max_tokens": 250}'
|
|
|
|
# 창의적 응답 (더 다양)
|
|
curl -X POST http://localhost:8000/api/v1/llm/ask \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "...", "temperature": 0.9, "max_tokens": 1000}'
|
|
```
|
|
|
|
---
|
|
|
|
## 성능 특성
|
|
|
|
| 작업 | 데이터셋 | 응답 시간 |
|
|
|------|---------|---------|
|
|
| LLM 쿼리 (캐시 미스) | - | 1-3초 |
|
|
| LLM 쿼리 (캐시 히트) | - | 50-100ms |
|
|
| 스트리밍 응답 (첫 토큰) | - | 500-800ms |
|
|
| RAG 메타데이터 | - | 100-300ms |
|
|
| 캐시 삭제 | 1K 키 | < 100ms |
|
|
| LLM 설정 변경 | - | < 50ms |
|
|
|
|
---
|
|
|
|
## 배포
|
|
|
|
### Docker
|
|
|
|
```dockerfile
|
|
FROM python:3.10-slim
|
|
|
|
WORKDIR /app
|
|
|
|
COPY requirements.txt .
|
|
RUN pip install -r requirements.txt
|
|
|
|
COPY . .
|
|
|
|
# Phase 7 앱 실행
|
|
CMD ["uvicorn", "ontology_platform.ont_platform.api.phase7_app:app", "--host", "0.0.0.0", "--port", "8000"]
|
|
```
|
|
|
|
### Kubernetes
|
|
|
|
```yaml
|
|
apiVersion: apps/v1
|
|
kind: Deployment
|
|
metadata:
|
|
name: ontology-phase7
|
|
spec:
|
|
replicas: 3
|
|
selector:
|
|
matchLabels:
|
|
app: ontology-phase7
|
|
template:
|
|
metadata:
|
|
labels:
|
|
app: ontology-phase7
|
|
spec:
|
|
containers:
|
|
- name: api
|
|
image: ontology-phase7:0.7.0
|
|
ports:
|
|
- containerPort: 8000
|
|
env:
|
|
- name: OPENAI_API_KEY
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: llm-secrets
|
|
key: openai-key
|
|
- name: NEO4J_URI
|
|
value: "bolt://neo4j:7687"
|
|
- name: REDIS_URL
|
|
value: "redis://redis:6379"
|
|
```
|
|
|
|
---
|
|
|
|
## 문제 해결
|
|
|
|
### Redis 연결 실패
|
|
|
|
```bash
|
|
# Redis 상태 확인
|
|
redis-cli ping
|
|
|
|
# Docker Redis 실행
|
|
docker run -d -p 6379:6379 redis:7.0
|
|
```
|
|
|
|
### LLM API 키 오류
|
|
|
|
```bash
|
|
# 환경 변수 확인
|
|
echo $OPENAI_API_KEY
|
|
|
|
# 유효한 API 키 설정
|
|
export OPENAI_API_KEY=sk-proj-xxx
|
|
```
|
|
|
|
### 높은 응답 시간
|
|
|
|
```bash
|
|
# 1. Redis 캐싱 활성화
|
|
# use_cache: true 설정
|
|
|
|
# 2. 토큰 제한 감소
|
|
# max_tokens: 250 설정
|
|
|
|
# 3. 온도 감소 (더 결정적)
|
|
# temperature: 0.3 설정
|
|
|
|
# 4. 로컬 LLM 사용 (프라이버시 + 속도)
|
|
# provider: local 설정
|
|
```
|
|
|
|
---
|
|
|
|
## 다음 단계
|
|
|
|
### Phase 8: 엔터프라이즈 기능
|
|
|
|
```
|
|
목표: 대규모 운영 지원
|
|
- 멀티테넌트 (여러 조직 동시 지원)
|
|
- 실시간 그래프 업데이트 (WebSocket)
|
|
- 변경 이력 추적 (감사 로그)
|
|
- 비용 관리 (API 호출당 요금)
|
|
- 고급 분석 (사용자별 통계)
|
|
```
|
|
|
|
---
|
|
|
|
## 정보
|
|
|
|
- **버전**: 0.7.0
|
|
- **마지막 업데이트**: 2026-05-14
|
|
- **지원 모델**: GPT-4, Claude 3, Llama 2, Mistral
|
|
- **캐시 TTL**: 1시간 (설정 가능)
|
|
|
|
---
|
|
|
|
**Phase 7 LLM 통합으로 지식 그래프를 기반으로 한 지능형 질문응답 시스템을 구축하세요!**
|