80d0ee4c5c7aabc00fa3cc556d9119be5f30ba8e
Ontology Platform
온톨로지 플랫폼은 웹에서 구조화된 지식(엔티티/관계)을 자동 추출, 검증, 저장하는 고속 시스템입니다.
Phase 0-4 전체 구현 완료 | 추출(10초) → 검증(<100ms) → 그래프 저장 → 벡터 검색
🚀 빠른 시작
1. 설치
# 기본 설치 (Phase 0-1: 추출)
pip install fastapi uvicorn pydantic trafilatura httpx
# Phase 2 추가 (동적 페이지)
pip install crawl4ai
# Phase 4 추가 (Neo4j)
pip install neo4j sentence-transformers
2. Phase 0-1만 사용 (가장 간단)
# API 서버 시작
python -m uvicorn ontology_platform.ont_platform.api.phase0_app:app --reload
# URL에서 추출
curl -X POST "http://localhost:8000/api/v1/extract/url?url=https://example.com"
3. Phase 4 (그래프 검색) 포함
# Neo4j 시작
docker-compose -f docker-compose.neo4j.yml up -d
# API 서버 시작
python -m uvicorn ontology_platform.ont_platform.api.phase0_app:app --reload
# 추출 → 수집 → 검색
curl -X POST "http://localhost:8000/api/v1/extract/url?url=https://example.com"
curl -X POST "http://localhost:8000/api/v1/search/ingest" -d '{"entities": [...], "relations": [...]}'
curl "http://localhost:8000/api/v1/search/vector?query=machine+learning"
📋 Phase별 기능
| Phase | 기능 | 시간 | 상태 |
|---|---|---|---|
| 0-1 | HTML 추출 (Trafilatura) | 10-15초 | ✅ |
| 2 | 동적 페이지 (Crawl4AI) | 20-30초 | ✅ |
| 3A | 경량 검증 (Pydantic) | <100ms | ✅ |
| 3B | SPARQL 검증 | <500ms | ✅ |
| 4 | Neo4j + 벡터 검색 | 50-200ms | ✅ |
🎯 사용 예시
예시 1: 기본 추출 (10초)
curl -X POST "http://localhost:8000/api/v1/extract/url?url=https://wikipedia.org/wiki/Python"
응답:
{
"url": "https://wikipedia.org/wiki/Python",
"title": "Python - Wikipedia",
"entities": [
{
"id": "E_1",
"label": "Python",
"type": "ProgrammingLanguage",
"confidence": 0.95
}
],
"relations": [...],
"extraction_time_sec": 9.5,
"validation_passed": true
}
예시 2: 동적 페이지 (25초)
curl -X POST "http://localhost:8000/api/v1/extract/url?url=https://app.example.com&profile=dynamic_page"
예시 3: 그래프 수집 + 검색
# 1. 추출
RESULT=$(curl -s -X POST "http://localhost:8000/api/v1/extract/url?url=https://example.com")
# 2. Neo4j에 수집
curl -X POST "http://localhost:8000/api/v1/search/ingest" \
-H "Content-Type: application/json" \
-d "{\"entities\": $(echo $RESULT | jq '.entities'), \"relations\": $(echo $RESULT | jq '.relations')}"
# 3. 벡터 검색
curl "http://localhost:8000/api/v1/search/vector?query=programming&limit=10"
# 4. 그래프 통계
curl "http://localhost:8000/api/v1/search/stats"
# 5. 엔티티 이웃
curl "http://localhost:8000/api/v1/search/entity/E_1?depth=1"
🔧 설정
Phase 선택 (validators.py)
# 경량 검증 (기본)
guard = OntologyGuard(validator_type="lightweight")
# SPARQL 검증
guard = OntologyGuard(validator_type="ontocast")
Neo4j 연결 (neo4j_adapter.py)
# 기본값
config = Neo4jConfig() # localhost:7687
# 커스텀
config = Neo4jConfig(
uri="bolt://custom-host:7687",
username="user",
password="pass",
database="mydb"
)
adapter = Neo4jAdapter(config=config)
📊 API 문서
서버 시작 후:
- Swagger UI: http://localhost:8000/docs
- ReDoc: http://localhost:8000/redoc
주요 엔드포인트
POST /api/v1/extract/url 추출
GET /api/v1/search/stats 통계
POST /api/v1/search/vector 벡터 검색
GET /api/v1/search/entity/{id} 이웃 탐색
POST /api/v1/search/ingest 그래프 수집
🧪 테스트
# Phase 0-1
python test_phase0_extraction.py
# Phase 2
python test_phase2_crawl.py
# Phase 3A
python test_phase3_validation.py
# Phase 3B
python test_phase3_option_b.py
# Phase 4
python test_phase4_integration.py
📦 의존성
- FastAPI: API 프레임워크
- Trafilatura: HTML 추출
- Crawl4AI: 동적 크롤링 (선택)
- Pydantic: 데이터 검증
- Neo4j: 그래프 DB (선택)
- SentenceTransformers: 벡터 임베딩 (선택)
🐳 Docker
# Neo4j만
docker-compose -f docker-compose.neo4j.yml up -d
# 전체 스택 (향후)
docker-compose up -d
📚 상세 문서
- 구현 요약 - Phase 0-4 전체 개요
- Phase 2 - Crawl4AI 동적 크롤링
- Phase 3A - 경량 검증
- Phase 3B - SPARQL 검증
- Phase 4 - Neo4j 그래프 + 벡터 검색
🎓 설계 원칙
- Phase-gated: 각 Phase는 선택사항
- Pluggable: 여러 검증 방식 지원
- Async: 높은 동시성
- Resilient: 의존성 부재 시에도 동작
💡 다음 단계
Phase 5: GraphRAG (선택)
- RDF ↔ Property Graph 변환
- Entity Resolver
- Subgraph retrieval
Advanced Features
- Critic loop (자동 수정)
- Few-shot learning
- Zero-shot 분류
🔗 관련 링크
📝 라이센스
MIT License
Version: 0.4.0 (Phase 0-4 완료)
Updated: 2026-05-14
Description
Languages
Jupyter Notebook
46%
Python
35.1%
TypeScript
12.4%
JavaScript
2.1%
CSS
0.9%
Other
3.2%