diff --git a/ontology_platform/.env.example b/ontology_platform/.env.example index f15f653..694e5dd 100644 --- a/ontology_platform/.env.example +++ b/ontology_platform/.env.example @@ -4,12 +4,34 @@ # ─── Phase 0: OntoCast Base ───────────────────────────────────────── ONTOCAST_WORKING_DIRECTORY=./data/working -# LLM Provider (OpenAI 또는 Ollama) +# ─── LLM Provider ─────────────────────────────────────────────────── +# 셋 중 하나를 골라 주석을 해제. ont_platform/config.py의 lenient 빌더가 +# OntoCast OpenAIModel enum 제약을 우회하므로, LM Studio처럼 임의 모델명도 +# LLM_MODEL_NAME에 그대로 적으면 된다. + +# [옵션 A] LM Studio (현재 기본) — http://127.0.0.1:1234 +# LM Studio의 "Developer > Reachable at" 주소를 그대로 LLM_BASE_URL에 입력. +# LLM_MODEL_NAME은 LM Studio가 로드한 모델 식별자와 일치시킨다. LLM_PROVIDER=openai -LLM_MODEL_NAME=gpt-4o-mini +LLM_MODEL_NAME=deepseek-r1-distill-qwen-7b +LLM_BASE_URL=http://127.0.0.1:1234/v1 +LLM_API_KEY=lm-studio LLM_TEMPERATURE=0.0 -LLM_API_KEY= -LLM_BASE_URL= + +# [옵션 B] Ollama 로컬 (https://ollama.com) +# 사전: `ollama pull qwen2.5` 등 +# LLM_PROVIDER=ollama +# LLM_MODEL_NAME=qwen2.5 +# LLM_BASE_URL=http://localhost:11434 +# LLM_API_KEY= +# LLM_TEMPERATURE=0.0 + +# [옵션 C] OpenAI 클라우드 +# LLM_PROVIDER=openai +# LLM_MODEL_NAME=gpt-4o-mini +# LLM_BASE_URL= +# LLM_API_KEY=sk-... +# LLM_TEMPERATURE=0.0 # ─── Server ───────────────────────────────────────────────────────── PORT=8000 diff --git a/ontology_platform/README.md b/ontology_platform/README.md index be0d9cb..718c565 100644 --- a/ontology_platform/README.md +++ b/ontology_platform/README.md @@ -38,7 +38,7 @@ docker compose up -d fuseki postgres redis # Phase 0~3 | 0.4 | ✅ 완료 | Robyn → FastAPI 재작성 (`/health`, `/info`, `/process`, `/flush`) | | 0.5 | ✅ 완료 | Pydantic Settings 정리 (filesystem 모드 강제) + 테스트 5개 | | 0.6 | ✅ 완료 | 통합 테스트(10개) + E2E 테스트(marker 분리) 작성 | -| 0.7 | ⚠️ 대기 | Acceptance Gate 0 — 실 환경에서 `pytest` 실행 필요 ([PHASE0_ACCEPTANCE_GATE.md](docs/phases/PHASE0_ACCEPTANCE_GATE.md)) | +| 0.7 | 🟡 부분완료 | Gate #2/#4 ✅ (unit 16/16, integration 10/10 통과). #1/#3 e2e 대기 ([PHASE0_ACCEPTANCE_GATE.md](docs/phases/PHASE0_ACCEPTANCE_GATE.md)) | **Next**: Acceptance Gate 0를 통과한 후 [PHASE1_NEXT_STEPS.md](docs/phases/PHASE1_NEXT_STEPS.md)로 진행. @@ -58,7 +58,7 @@ ontology_platform/ │ └── phases/ ← Phase별 작업 로그 ├── vendored/ │ └── ontocast/ ← Phase 0.1에서 추가 -├── platform/ ← 우리가 작성하는 코드 +├── ont_platform/ ← 우리가 작성하는 코드 (※ Python 내장 `platform` 모듈과 이름 충돌을 피하기 위해 `platform/`에서 변경됨) │ ├── api/ ← Phase 0.4 (FastAPI) │ ├── core/ │ │ ├── extractors/ ← Phase 1 (Trafilatura) diff --git a/ontology_platform/docs/phases/PHASE0_ACCEPTANCE_GATE.md b/ontology_platform/docs/phases/PHASE0_ACCEPTANCE_GATE.md index 2a37918..73fae7a 100644 --- a/ontology_platform/docs/phases/PHASE0_ACCEPTANCE_GATE.md +++ b/ontology_platform/docs/phases/PHASE0_ACCEPTANCE_GATE.md @@ -6,16 +6,19 @@ | # | Acceptance Gate 항목 | 상태 | 검증 방법 | |---|---|---|---| -| 1 | 단일 PDF/JSON 입력 → ontology TTL + facts TTL이 filesystem에 생성됨 | ⚠️ **코드 준비 완료, 실행 검증 보류** | `tests/e2e/test_phase0_full_pipeline.py`가 검증하나 LLM_API_KEY/Python 환경 필요 | -| 2 | `/health`, `/info`, `/process` (FastAPI) 정상 동작 | ✅ **코드 작성 + 통합 테스트 통과 예상** | `tests/integration/test_api_smoke.py` 11개 케이스 | -| 3 | BudgetTracker가 LLM call/triple count를 정확히 기록 | ⚠️ **코드 준비 완료, 실 LLM 호출 검증 보류** | 통합 테스트는 mock 검증, e2e 테스트가 실제 검증 | +| 1 | 단일 PDF/JSON 입력 → ontology TTL + facts TTL이 filesystem에 생성됨 | ⚠️ **e2e 검증 대기** (로컬 LLM/API 키 필요) | `tests/e2e/test_phase0_full_pipeline.py` | +| 2 | `/health`, `/info`, `/process` (FastAPI) 정상 동작 | ✅ **통합 테스트 10/10 통과** (2026-05-14) | `tests/integration/test_api_smoke.py` | +| 3 | BudgetTracker가 LLM call/triple count를 정확히 기록 | ⚠️ **e2e 검증 대기** (mock 검증은 통합 테스트로 통과) | e2e 테스트가 실제 검증 | | 4 | LangGraph 워크플로우 (CONVERT→CHUNK→...→SERIALIZE) 전 노드 traceable | ✅ **OntoCast 원본 워크플로우 무수정 채택** | `vendored/ontocast/ontocast/stategraph/` 그대로 사용 | -⚠️ **현재 환경에서 자동 실행이 안 되는 이유**: -1. 시스템에 Python 인터프리터가 설치되어 있지 않음 (`python.exe`가 Microsoft Store 별칭만 있음, `py` 없음) -2. LLM API 키가 환경변수에 없음 +추가로 **단위 테스트 16/16 통과** (test_convert_document 7, test_platform_config 5, test_select_ontology 4). -따라서 **다음 작업자(또는 운영 환경)에서 아래 절차를 한 번 실행하여 4개 체크박스를 모두 통과 처리해야 한다**. 코드는 준비 완료. +**현재 진척 (2026-05-14)**: +- Python 3.13.13 환경 + `pip install -e ".[dev]"` 완료 +- `pip install -e vendored/ontocast` 로 OntoCast 의존성 설치 완료 +- 패키지 이름 충돌 수정: `platform/` → `ont_platform/` (Python 내장 `platform` 모듈과 충돌) +- 단위 + 통합 테스트 26/26 모두 통과 +- **남은 작업**: e2e 테스트 (Acceptance Gate #1, #3) 실행 — 로컬 Ollama 또는 OpenAI 키 필요 ## 다음 작업자가 실행할 검증 절차 @@ -54,8 +57,20 @@ pytest tests/unit tests/integration -v ### 3) End-to-end 검증 (Acceptance Gate #1, #3, #4) +LLM 호출이 실제로 일어남. OpenAI는 비용 발생, Ollama는 로컬에서 무료. + ```powershell -# LLM 호출이 일어남. 실 비용 발생. +# (A) Ollama 로컬 사용 (권장 — 비용 무료) +# 사전: Ollama 설치 후 `ollama pull qwen2.5` +$env:LLM_PROVIDER = "ollama" +$env:LLM_MODEL_NAME = "qwen2.5" +$env:LLM_BASE_URL = "http://localhost:11434" +pytest tests/e2e -m e2e -v + +# (B) OpenAI 사용 +$env:LLM_PROVIDER = "openai" +$env:LLM_MODEL_NAME = "gpt-4o-mini" +$env:LLM_API_KEY = "sk-..." pytest tests/e2e -m e2e -v ``` @@ -70,7 +85,7 @@ pytest tests/e2e -m e2e -v ```powershell # 서버 기동 -uvicorn platform.api.main:app --reload +uvicorn ont_platform.api.main:app --reload # 다른 셸에서 curl http://localhost:8000/health @@ -100,4 +115,5 @@ curl -X POST http://localhost:8000/process ` | 일자 | 검증자 | 결과 | |---|---|---| | 2026-05-13 | (코드 작성: ontology-platform agent) | 코드 준비 완료. 실 환경 검증 보류. | +| 2026-05-14 | lasta + Claude | **unit 16/16, integration 10/10 통과** (Gate #2 ✅). 패키지 이름 충돌 수정 (`platform`→`ont_platform`). e2e는 LLM 필요로 대기. | | ____-__-__ | ________________ | __________________________________ | diff --git a/ontology_platform/platform/__init__.py b/ontology_platform/ont_platform/__init__.py similarity index 100% rename from ontology_platform/platform/__init__.py rename to ontology_platform/ont_platform/__init__.py diff --git a/ontology_platform/platform/api/__init__.py b/ontology_platform/ont_platform/api/__init__.py similarity index 100% rename from ontology_platform/platform/api/__init__.py rename to ontology_platform/ont_platform/api/__init__.py diff --git a/ontology_platform/platform/api/deps.py b/ontology_platform/ont_platform/api/deps.py similarity index 90% rename from ontology_platform/platform/api/deps.py rename to ontology_platform/ont_platform/api/deps.py index ec683e4..af24cbd 100644 --- a/ontology_platform/platform/api/deps.py +++ b/ontology_platform/ont_platform/api/deps.py @@ -35,7 +35,7 @@ from ontocast.toolbox import ToolBox # noqa: E402 import importlib -platform_config = importlib.import_module("platform.config") +platform_config = importlib.import_module("ont_platform.config") logger = logging.getLogger(__name__) @@ -67,7 +67,10 @@ async def initialize_app_context( settings = settings or platform_config.load_settings() ontocast_config = platform_config.build_ontocast_config(settings) - tools = ToolBox(ontocast_config) + # ToolBox.__init__ 내부에서 LLMTool.create()가 `asyncio.run()`을 호출한다. + # FastAPI lifespan/테스트가 이미 async 컨텍스트면 이중 loop 충돌이 나므로, + # 별도 스레드에서 sync 생성자를 실행한다. + tools = await asyncio.to_thread(ToolBox, ontocast_config) # OntoCast's ToolBox.initialize is async; do it here so a request doesn't # have to pay the cost. await tools.initialize() diff --git a/ontology_platform/platform/api/main.py b/ontology_platform/ont_platform/api/main.py similarity index 98% rename from ontology_platform/platform/api/main.py rename to ontology_platform/ont_platform/api/main.py index f5ec0da..9a71078 100644 --- a/ontology_platform/platform/api/main.py +++ b/ontology_platform/ont_platform/api/main.py @@ -38,14 +38,14 @@ if str(_VENDORED_ONTOCAST) not in sys.path: from ontocast.onto.enum import RenderMode # noqa: E402 from ontocast.onto.state import AgentState # noqa: E402 -from platform.api.deps import ( # noqa: E402 +from ont_platform.api.deps import ( # noqa: E402 AppContext, RunnableConfig, get_app_context, initialize_app_context, ) -platform_config = importlib.import_module("platform.config") +platform_config = importlib.import_module("ont_platform.config") logger = logging.getLogger(__name__) @@ -381,7 +381,7 @@ app = create_app() @click.option("--reload", is_flag=True, default=False) def cli(host: str, port: int, reload: bool) -> None: # noqa: FBT001 """Console entry point: ``ontology-platform`` (see pyproject.toml).""" - uvicorn.run("platform.api.main:app", host=host, port=port, reload=reload) + uvicorn.run("ont_platform.api.main:app", host=host, port=port, reload=reload) if __name__ == "__main__": diff --git a/ontology_platform/platform/config.py b/ontology_platform/ont_platform/config.py similarity index 82% rename from ontology_platform/platform/config.py rename to ontology_platform/ont_platform/config.py index 85c1653..44f652f 100644 --- a/ontology_platform/platform/config.py +++ b/ontology_platform/ont_platform/config.py @@ -16,6 +16,7 @@ build a ``Config`` instance and pass it to ``ToolBox``. from __future__ import annotations import logging +import os import sys from enum import IntEnum from pathlib import Path @@ -138,6 +139,31 @@ class PlatformSettings(BaseSettings): return self +def _build_llm_config_lenient() -> LLMConfig: + """env vars로부터 LLMConfig 생성. OntoCast의 OpenAIModel enum validation을 우회한다. + + Why: LM Studio / vLLM / llama.cpp 등 OpenAI-호환 로컬 서버는 임의의 model + identifier를 쓰며(예: `deepseek-r1-distill-qwen-7b`), 이는 OntoCast의 정해진 + enum(`gpt-4o`, `gpt-4o-mini`, ...)에 들어가지 않는다. ChatOpenAI는 model을 + 문자열로 받으므로 enum 강제만 풀면 OntoCast 다른 코드 경로는 그대로 동작한다. + + `LLMConfig.model_construct`는 Pydantic V2의 validation 우회 생성자다. + """ + provider_raw = (os.getenv("LLM_PROVIDER") or "openai").lower() + model_name = os.getenv("LLM_MODEL_NAME") or "gpt-4o-mini" + temperature_raw = os.getenv("LLM_TEMPERATURE") or "0.0" + base_url = os.getenv("LLM_BASE_URL") or None + api_key = os.getenv("LLM_API_KEY") or None + + return LLMConfig.model_construct( + provider=provider_raw, + model_name=model_name, + temperature=float(temperature_raw), + base_url=base_url, + api_key=api_key, + ) + + def _empty_neo4j_config() -> Neo4jConfig: """A Neo4jConfig with no URI/auth so ToolBox skips Neo4j initialization. @@ -171,9 +197,18 @@ def build_ontocast_config(settings: PlatformSettings) -> OntoCastConfig: - Neo4j and Fuseki are forcibly disabled regardless of NEO4J_*/FUSEKI_* env vars in the shell. They will be wired in Phase 4. """ - # Start from defaults that pull in any LLM_*/CHUNK_*/AGG_* env vars - # via each section's own SettingsConfigDict. - tool_cfg = ToolConfig() + # OntoCast의 OpenAIModel enum은 클라우드 모델만 허용한다. LM Studio 등 + # 임의의 모델명을 쓰는 로컬 서버는 ToolConfig() 생성 단계에서 검증이 실패 + # 하므로, ToolConfig를 만들 동안만 LLM_MODEL_NAME을 비우고 lenient 빌더로 + # 교체한다. CHUNK_*/AGG_* 등 다른 섹션 env는 그대로 흘러가도록 유지한다. + saved_model = os.environ.pop("LLM_MODEL_NAME", None) + try: + tool_cfg = ToolConfig() + finally: + if saved_model is not None: + os.environ["LLM_MODEL_NAME"] = saved_model + + tool_cfg.llm_config = _build_llm_config_lenient() # Override paths from the platform settings. tool_cfg.path_config = PathConfig( diff --git a/ontology_platform/platform/core/__init__.py b/ontology_platform/ont_platform/core/__init__.py similarity index 100% rename from ontology_platform/platform/core/__init__.py rename to ontology_platform/ont_platform/core/__init__.py diff --git a/ontology_platform/platform/core/crawler/__init__.py b/ontology_platform/ont_platform/core/crawler/__init__.py similarity index 100% rename from ontology_platform/platform/core/crawler/__init__.py rename to ontology_platform/ont_platform/core/crawler/__init__.py diff --git a/ontology_platform/platform/core/extractors/__init__.py b/ontology_platform/ont_platform/core/extractors/__init__.py similarity index 100% rename from ontology_platform/platform/core/extractors/__init__.py rename to ontology_platform/ont_platform/core/extractors/__init__.py diff --git a/ontology_platform/platform/core/projection/__init__.py b/ontology_platform/ont_platform/core/projection/__init__.py similarity index 100% rename from ontology_platform/platform/core/projection/__init__.py rename to ontology_platform/ont_platform/core/projection/__init__.py diff --git a/ontology_platform/platform/core/validation/__init__.py b/ontology_platform/ont_platform/core/validation/__init__.py similarity index 100% rename from ontology_platform/platform/core/validation/__init__.py rename to ontology_platform/ont_platform/core/validation/__init__.py diff --git a/ontology_platform/platform/storage/__init__.py b/ontology_platform/ont_platform/storage/__init__.py similarity index 100% rename from ontology_platform/platform/storage/__init__.py rename to ontology_platform/ont_platform/storage/__init__.py diff --git a/ontology_platform/platform/workflow/__init__.py b/ontology_platform/ont_platform/workflow/__init__.py similarity index 100% rename from ontology_platform/platform/workflow/__init__.py rename to ontology_platform/ont_platform/workflow/__init__.py diff --git a/ontology_platform/pyproject.toml b/ontology_platform/pyproject.toml index 67c086c..1d87b0b 100644 --- a/ontology_platform/pyproject.toml +++ b/ontology_platform/pyproject.toml @@ -113,16 +113,16 @@ dev = [ ] [project.scripts] -ontology-platform = "platform.api.main:cli" +ontology-platform = "ont_platform.api.main:cli" [tool.hatch.build.targets.wheel] -packages = ["platform"] +packages = ["ont_platform"] # ─── Ruff (linter + formatter) ──────────────────────────────────────── [tool.ruff] line-length = 100 target-version = "py312" -src = ["platform", "tests"] +src = ["ont_platform", "tests"] extend-exclude = ["vendored"] # vendored OntoCast 등은 원본 유지 [tool.ruff.lint] diff --git a/ontology_platform/tests/e2e/conftest.py b/ontology_platform/tests/e2e/conftest.py index d66aa48..f810e75 100644 --- a/ontology_platform/tests/e2e/conftest.py +++ b/ontology_platform/tests/e2e/conftest.py @@ -1,12 +1,17 @@ """E2E test configuration. -These tests are SKIPPED by default. To run them set ``LLM_API_KEY`` and -``LLM_PROVIDER`` (and any model overrides) in the environment, then run: +These tests are SKIPPED by default unless a valid LLM provider is configured. +The repo's `.env` file is auto-loaded so the same settings used by the app +also drive the test run. +Provider별 통과 조건: +- openai : `LLM_API_KEY` 필요 (LM Studio 등 OpenAI-호환 로컬 서버는 더미 키도 OK) +- ollama : 별도 키 불필요 (로컬 데몬만 동작하면 됨) + +To run: pytest tests/e2e -m e2e -They exercise the full OntoCast workflow end-to-end — LLM calls included — -and so they cost real money. Keep them out of CI default runs. +LLM 호출이 실제로 일어남. OpenAI 클라우드는 비용 발생, 로컬 LLM은 무료. """ from __future__ import annotations @@ -22,18 +27,47 @@ for p in (REPO_ROOT, REPO_ROOT / "vendored" / "ontocast"): if str(p) not in sys.path: sys.path.insert(0, str(p)) +# Auto-load .env so e2e tests pick up the same LLM settings as the app. +_env_path = REPO_ROOT / ".env" +if _env_path.exists(): + try: + from dotenv import load_dotenv + + load_dotenv(_env_path, override=False) + except ImportError: + # python-dotenv가 없으면 직접 간단 파싱. + for line in _env_path.read_text(encoding="utf-8").splitlines(): + line = line.strip() + if not line or line.startswith("#") or "=" not in line: + continue + key, _, value = line.partition("=") + os.environ.setdefault(key.strip(), value.strip()) + + +def _llm_configured() -> bool: + """Provider별로 e2e 실행 가능 여부 판단.""" + provider = (os.environ.get("LLM_PROVIDER") or "openai").lower() + if provider == "ollama": + # Ollama는 별도 API 키 불필요. + return True + # openai 또는 OpenAI-호환 로컬 서버 — 더미 키라도 들어 있어야 OK. + return bool(os.environ.get("LLM_API_KEY")) + def pytest_collection_modifyitems( config: pytest.Config, items: list[pytest.Item] ) -> None: - """Skip all e2e tests when LLM_API_KEY is not configured.""" - if os.environ.get("LLM_API_KEY"): + """Skip e2e tests unless the LLM provider is properly configured.""" + if _llm_configured(): return + provider = (os.environ.get("LLM_PROVIDER") or "openai").lower() skip_marker = pytest.mark.skip( - reason="E2E tests require LLM_API_KEY in environment." + reason=( + f"E2E tests require LLM provider config. " + f"Provider={provider!r}, LLM_API_KEY={'set' if os.environ.get('LLM_API_KEY') else 'missing'}." + ) ) for item in items: - # Only apply to tests in this directory tree. if "tests/e2e" in str(item.fspath).replace("\\", "/"): item.add_marker(skip_marker) diff --git a/ontology_platform/tests/e2e/test_phase0_full_pipeline.py b/ontology_platform/tests/e2e/test_phase0_full_pipeline.py index 9af2249..4626b26 100644 --- a/ontology_platform/tests/e2e/test_phase0_full_pipeline.py +++ b/ontology_platform/tests/e2e/test_phase0_full_pipeline.py @@ -26,9 +26,9 @@ REPO_ROOT = Path(__file__).resolve().parents[2] if str(REPO_ROOT) not in sys.path: sys.path.insert(0, str(REPO_ROOT)) -main_module = importlib.import_module("platform.api.main") -deps_module = importlib.import_module("platform.api.deps") -platform_config = importlib.import_module("platform.config") +main_module = importlib.import_module("ont_platform.api.main") +deps_module = importlib.import_module("ont_platform.api.deps") +platform_config = importlib.import_module("ont_platform.config") @pytest.mark.e2e @@ -40,7 +40,12 @@ async def test_full_pipeline_writes_ontology_and_facts( # leak between runs. monkeypatch.setenv("ONTOCAST_WORKING_DIRECTORY", str(tmp_path / "work")) # Honor whatever LLM provider the operator configured. - assert os.environ.get("LLM_API_KEY"), "LLM_API_KEY must be set for e2e" + # OpenAI는 API 키 필수, Ollama는 로컬 데몬만 떠 있으면 됨. + provider = os.environ.get("LLM_PROVIDER", "openai").lower() + if provider == "openai": + assert os.environ.get("LLM_API_KEY"), ( + "LLM_API_KEY must be set for e2e with OpenAI provider" + ) # Force a fresh AppContext so the new working_directory wins. deps_module.reset_app_context_for_testing() @@ -49,11 +54,28 @@ async def test_full_pipeline_writes_ontology_and_facts( app = main_module.create_app() - # Tiny but ontology-rich payload. + # OntoCast SemanticChunker가 HDBSCAN + UMAP을 쓰므로 최소 문장 수가 + # 필요하다. 2문장짜리 toy 입력은 "k must be ≤ training points"로 실패한다. + # Acceptance Gate #1은 "단일 PDF/JSON → TTL 생성"이지 문장 수와 무관하므로, + # ontology-rich한 짧은 단락을 충분한 문장 수로 늘려준다. payload = { "text": ( - "Alice works at Acme Corporation in Berlin. " - "Acme Corporation manufactures bicycles." + "Alice Carter works at Acme Corporation in Berlin. " + "Acme Corporation is a manufacturing company founded in 1992. " + "Acme manufactures bicycles and electric scooters. " + "Bob Lee is the chief engineer at Acme Corporation. " + "He reports to Alice Carter, who heads the engineering division. " + "Acme's main factory is located in Berlin, Germany. " + "The company also operates a research center in Munich. " + "Carol Schmidt leads research at the Munich center. " + "She previously worked at Globex Industries in Hamburg. " + "Globex Industries is a competitor in the bicycle market. " + "Acme exports bicycles to France, Italy, and Spain. " + "The product line includes road bikes, mountain bikes, and city bikes. " + "Alice Carter graduated from the Technical University of Berlin. " + "Bob Lee holds a doctorate in mechanical engineering. " + "Acme employs around 450 people across its three sites. " + "The company reported annual revenue of 120 million euros last year." ), "ontology_user_instruction": "Focus on person-organization-location relations.", "facts_user_instruction": "Extract employment and manufacturing facts.", diff --git a/ontology_platform/tests/integration/conftest.py b/ontology_platform/tests/integration/conftest.py index 6fc8729..c67e7cc 100644 --- a/ontology_platform/tests/integration/conftest.py +++ b/ontology_platform/tests/integration/conftest.py @@ -1,7 +1,7 @@ """Shared fixtures for integration tests. -Sets up the Python path so `import platform.api.main` resolves the local -package (not the stdlib `platform` module) and exposes helpers that turn +Sets up the Python path so `import ont_platform.api.main` resolves the local +package and exposes helpers that turn the FastAPI app into a controllable test harness. """ diff --git a/ontology_platform/tests/integration/test_api_smoke.py b/ontology_platform/tests/integration/test_api_smoke.py index 98620f1..a612794 100644 --- a/ontology_platform/tests/integration/test_api_smoke.py +++ b/ontology_platform/tests/integration/test_api_smoke.py @@ -18,6 +18,7 @@ from __future__ import annotations import importlib import json import sys +from contextlib import asynccontextmanager from pathlib import Path from types import SimpleNamespace from typing import Any @@ -33,8 +34,8 @@ VENDORED_ONTOCAST = REPO_ROOT / "vendored" / "ontocast" if str(VENDORED_ONTOCAST) not in sys.path: sys.path.insert(0, str(VENDORED_ONTOCAST)) -main_module = importlib.import_module("platform.api.main") -deps_module = importlib.import_module("platform.api.deps") +main_module = importlib.import_module("ont_platform.api.main") +deps_module = importlib.import_module("ont_platform.api.deps") from ontocast.onto.enum import RenderMode # noqa: E402 @@ -84,6 +85,11 @@ def _make_mock_context(workflow_chunks: list[dict[str, Any]]) -> SimpleNamespace ) +@asynccontextmanager +async def _noop_lifespan(app): # noqa: ARG001 + yield + + def _client_with_context(ctx: SimpleNamespace) -> TestClient: """Return a TestClient whose ``get_app_context`` returns the given ctx. @@ -92,9 +98,8 @@ def _client_with_context(ctx: SimpleNamespace) -> TestClient: surface in test output. """ app = main_module.create_app() + app.router.lifespan_context = _noop_lifespan app.dependency_overrides[deps_module.get_app_context] = lambda: ctx - # `TestClient` runs the lifespan by default; disable it because we're - # providing the context manually. return TestClient(app, raise_server_exceptions=True, backend="asyncio") diff --git a/ontology_platform/tests/unit/test_convert_document.py b/ontology_platform/tests/unit/test_convert_document.py index 73d8188..9a6ae8c 100644 --- a/ontology_platform/tests/unit/test_convert_document.py +++ b/ontology_platform/tests/unit/test_convert_document.py @@ -23,7 +23,11 @@ VENDORED_ONTOCAST = REPO_ROOT / "vendored" / "ontocast" if str(VENDORED_ONTOCAST) not in sys.path: sys.path.insert(0, str(VENDORED_ONTOCAST)) -from ontocast.agent import convert_document as convert_document_module # noqa: E402 +import sys + +import ontocast.agent # __init__.py 실행으로 서브모듈이 sys.modules에 등록됨 # noqa: E402 + +convert_document_module = sys.modules["ontocast.agent.convert_document"] from ontocast.onto.enum import Status # noqa: E402 diff --git a/ontology_platform/tests/unit/test_platform_config.py b/ontology_platform/tests/unit/test_platform_config.py index c568e7e..6fd3711 100644 --- a/ontology_platform/tests/unit/test_platform_config.py +++ b/ontology_platform/tests/unit/test_platform_config.py @@ -24,7 +24,7 @@ if str(PLATFORM_ROOT) not in sys.path: # by relying on the package being on sys.path before site-packages.) import importlib -platform_config = importlib.import_module("platform.config") +platform_config = importlib.import_module("ont_platform.config") def _clear_settings_env(monkeypatch: pytest.MonkeyPatch) -> None: diff --git a/ontology_platform/tests/unit/test_select_ontology.py b/ontology_platform/tests/unit/test_select_ontology.py index b7065c3..6625388 100644 --- a/ontology_platform/tests/unit/test_select_ontology.py +++ b/ontology_platform/tests/unit/test_select_ontology.py @@ -29,7 +29,11 @@ if str(VENDORED_ONTOCAST) not in sys.path: sys.path.insert(0, str(VENDORED_ONTOCAST)) # Imports must come after sys.path tweak. -from ontocast.agent import select_ontology as select_ontology_module # noqa: E402 +import sys + +import ontocast.agent # __init__.py 실행으로 서브모듈이 sys.modules에 등록됨 # noqa: E402 + +select_ontology_module = sys.modules["ontocast.agent.select_ontology"] from ontocast.onto.enum import Status # noqa: E402 from ontocast.onto.null import NULL_ONTOLOGY # noqa: E402 diff --git a/ontology_platform/vendored/ontocast/VENDORED_MODIFICATIONS.md b/ontology_platform/vendored/ontocast/VENDORED_MODIFICATIONS.md index 031ee95..e588b71 100644 --- a/ontology_platform/vendored/ontocast/VENDORED_MODIFICATIONS.md +++ b/ontology_platform/vendored/ontocast/VENDORED_MODIFICATIONS.md @@ -28,6 +28,20 @@ - **회귀 테스트**: `tests/unit/test_convert_document.py` (7 케이스 — 단일 PDF / 단일 JSON / 다중 PDF / 다중 JSON first-wins / 미지원 확장자 / 빈 입력 / 혼합 PDF+JSON) - **근거**: docs/통합설계서.md §5 Phase 0, OntoCast 분석 §13.1 / §21.1 (#4) +### 2026-05-14 — node_factories.py: BudgetTracker가 LLM 호출을 기록하지 못하던 버그 수정 +- **파일**: `ontocast/stategraph/node_factories.py` +- **문제**: `make_render_ontology_node`(line 91)와 `make_render_facts_node`(line 256)에서 병렬 unit 처리 시 `base_state = state.model_copy(deep=True)`로 root state를 deep-copy 한 뒤 `budget_tracker=base_state.budget_tracker`를 child state(`UnitOntologyState`/`UnitFactsState`)에 전달했다. 그 결과 LLM이 add_usage()를 호출해도 *복사본* 인스턴스만 갱신되고 root state.budget_tracker는 영원히 0인 채로 남아 BudgetTracker가 LLM 호출 수, chars sent/received를 전혀 기록하지 못했다. e2e 실행 시 LM Studio 로그에 LLM 호출이 분명히 들어왔는데(token count 검증됨) workflow_state["budget_tracker"].calls_count == 0으로 응답되는 증상으로 확인됨. +- **수정**: 두 곳 모두 `base_state.budget_tracker` → `state.budget_tracker`로 변경. `base_state = state.model_copy(deep=True)` 라인 자체가 budget_tracker 추출에만 쓰였으므로 함께 제거. Python int `+= 1` 연산은 GIL 하에서 원자적이라 병렬 process_unit 간 race condition은 무시 가능한 수준. +- **회귀 테스트**: e2e (`tests/e2e/test_phase0_full_pipeline.py`)에서 `budget.calls_count > 0` 검증. +- **근거**: docs/phases/PHASE0_ACCEPTANCE_GATE.md (Gate #3 통과를 위한 차단 이슈). + +### 2026-05-14 — render_ontology.py: Bootstrap KeyError 수정 +- **파일**: `ontocast/agent/render_ontology.py` +- **문제**: `vendored/ontocast/ontocast/prompt/render_ontology.py:44`의 `general_ontology_instruction = f"""...{prefix_instruction}..."""`이 f-string이라 모듈 로드 시점에 `prefix_instruction` 변수가 즉시 치환됨. 그 결과 최종 문자열에는 `prefix_instruction` 자체가 가진 `{ontology_prefix}` placeholder만 남는다. `render_ontology_update`는 호출 시 `ontology_prefix=current.prefix`도 함께 넘기지만 `render_ontology_fresh`는 누락 → Bootstrap 단계에서 `KeyError: 'ontology_prefix'`. +- **수정**: `render_ontology_fresh()`의 `.format()` 호출에 `ontology_prefix=""`를 추가. Bootstrap 시점에는 prefix가 아직 LLM에 의해 정의되지 않으므로 빈 문자열이 의미적으로 올바름. +- **회귀 테스트**: e2e (`tests/e2e/test_phase0_full_pipeline.py`)에서 검증. +- **근거**: docs/phases/PHASE0_ACCEPTANCE_GATE.md (Gate #1 통과를 위한 차단 이슈). + ### 2026-05-13 — Phase 0.2: select_ontology.py None-index 버그 수정 - **파일**: `ontocast/agent/select_ontology.py` - **문제**: dynamic Pydantic 모델은 `answer_index ∈ [1, num_ontologies + 1]`을 강제하지만 코드는 `answer_index == 0`을 "None"으로 처리. 0은 Pydantic 검증을 통과할 수 없어 dead code였으며, LLM이 None을 선택할 때마다(`num_ontologies + 1` 반환) defensive branch로 빠져 WARNING 로그가 찍혔다. diff --git a/ontology_platform/vendored/ontocast/ontocast/agent/render_ontology.py b/ontology_platform/vendored/ontocast/ontocast/agent/render_ontology.py index 56c66ba..1b434b4 100644 --- a/ontology_platform/vendored/ontocast/ontocast/agent/render_ontology.py +++ b/ontology_platform/vendored/ontocast/ontocast/agent/render_ontology.py @@ -5,6 +5,11 @@ human-readable formats, making the ontological knowledge more accessible and understandable. The agent decides between generating bare Turtle for fresh ontologies and SPARQL operations for updates. +# MODIFIED 2026-05-14 (ontology_platform): render_ontology_fresh()의 .format() +# 호출에 ontology_prefix 키가 누락되어 Bootstrap 단계에서 KeyError 발생. +# general_ontology_instruction은 f-string 모듈 로드 시점에 prefix_instruction이 +# 치환되며 그 안의 {ontology_prefix} 플레이스홀더만 남는다. Fresh 시점에는 +# prefix가 아직 LLM에 의해 정해지지 않으므로 빈 문자열을 전달한다. """ import logging @@ -109,8 +114,13 @@ async def render_ontology_fresh( output_instruction = output_instruction_ttl ontology_ttl = "" improvement_instruction_str = "" + # MODIFIED 2026-05-14 (ontology_platform): Bootstrap에서는 prefix가 아직 + # 정해지지 않았으므로 ontology_prefix를 빈 문자열로 채워 KeyError를 피한다. + # prefix_instruction= 인자는 f-string으로 이미 박혔으므로 무시되지만 호환을 + # 위해 그대로 둔다. general_ontology_instruction_str = general_ontology_instruction.format( - prefix_instruction=prefix_instruction_fresh + prefix_instruction=prefix_instruction_fresh, + ontology_prefix="", ) text_chapter = text_template.format(text=state.content_unit.text) diff --git a/ontology_platform/vendored/ontocast/ontocast/stategraph/node_factories.py b/ontology_platform/vendored/ontocast/ontocast/stategraph/node_factories.py index d228a05..02ab233 100644 --- a/ontology_platform/vendored/ontocast/ontocast/stategraph/node_factories.py +++ b/ontology_platform/vendored/ontocast/ontocast/stategraph/node_factories.py @@ -1,3 +1,7 @@ +# MODIFIED 2026-05-14 (ontology_platform): make_render_ontology_node와 +# make_render_facts_node에서 `state.model_copy(deep=True)`로 budget_tracker가 +# deep-copy되어 root state의 BudgetTracker가 0인 채로 남던 버그 수정. +# 자세한 내역은 vendored/ontocast/VENDORED_MODIFICATIONS.md 참고. import asyncio import logging @@ -88,12 +92,19 @@ def make_render_ontology_node(tools: ToolBox): async def process_unit(unit_index: int) -> tuple[int, UnitOntologyState]: async with semaphore: - base_state = state.model_copy(deep=True) + # MODIFIED 2026-05-14 (ontology_platform): 원래 코드는 + # `state.model_copy(deep=True)`로 budget_tracker를 deep-copy해 + # child state에 넘겼다. 그 결과 LLM이 add_usage()를 호출해도 + # *복사본* 인스턴스만 갱신되고 root state.budget_tracker는 + # 영원히 0인 채로 남아 BudgetTracker가 LLM 호출을 기록하지 + # 못했다. 원본 인스턴스를 공유해 add_usage가 root state까지 + # 직접 반영되도록 수정. (Python int += 1은 GIL 하에서 원자적이라 + # 병렬 process_unit 간 race-condition 위험은 무시 가능) ontology_state = UnitOntologyState( content_unit=state.content_units[unit_index], ontology_snapshot=state.current_ontology, ontology_user_instruction=state.ontology_user_instruction, - budget_tracker=base_state.budget_tracker, + budget_tracker=state.budget_tracker, max_visits_per_node=tools.config.server.max_visits_per_node, current_domain=state.current_domain, ontology_max_triples=tools.config.server.ontology_max_triples, @@ -253,12 +264,14 @@ def make_render_facts_node(tools: ToolBox): async def process_unit(unit_index: int) -> tuple[int, UnitFactsState]: async with semaphore: - base_state = state.model_copy(deep=True) + # MODIFIED 2026-05-14 (ontology_platform): same fix as + # make_render_ontology_node — share the root budget_tracker + # instead of a deep-copied detached instance. facts_state = UnitFactsState( content_unit=state.content_units[unit_index], ontology_snapshot=state.current_ontology, facts_user_instruction=state.facts_user_instruction, - budget_tracker=base_state.budget_tracker, + budget_tracker=state.budget_tracker, max_visits_per_node=tools.config.server.max_visits_per_node, ) result = await facts_loop(facts_state, atomic_tools)