Files
AI/ontology_platform/tests/unit/test_convert_document.py
lasta ec4f9a64f6 Phase 0.7 — Acceptance Gate 자동화 + LM Studio 통합 + OntoCast 버그 수정
- platform/ → ont_platform/ rename
  Python 내장 platform 모듈과 이름 충돌. numpy/scipy가 platform.machine() 호출 시
  우리 패키지를 가져와 AttributeError. ont_platform으로 변경하고 pyproject.toml,
  ont_platform/**, tests/** import 경로 모두 업데이트.

- ont_platform/config.py: lenient LLM builder 추가
  LM Studio/vLLM 등 OpenAI-호환 로컬 서버가 임의 모델 식별자(예: deepseek-r1-distill-
  qwen-7b)를 쓸 수 있도록 OntoCast의 OpenAIModel enum validation을 Pydantic
  model_construct로 우회. ToolConfig() 생성 시 충돌을 막기 위해 LLM_MODEL_NAME을
  잠시 비웠다가 lenient 인스턴스로 교체.

- ont_platform/api/deps.py: ToolBox 초기화를 asyncio.to_thread로 격리
  LLMTool.create()가 내부에서 asyncio.run()을 부르는데 lifespan/테스트가 이미
  async 컨텍스트라 이중 loop 충돌. 별도 스레드에서 sync 생성자 실행.

- 테스트 인프라 정비
  * tests/integration/test_api_smoke.py: TestClient 구버전 starlette 호환을 위해
    lifespan='off' 대신 app.router.lifespan_context = noop 패턴 적용.
  * tests/unit/test_convert_document.py, test_select_ontology.py: ontocast.agent
    __init__.py가 re-export한 함수가 서브모듈을 가리는 문제로 sys.modules에서
    실제 모듈 객체 직접 추출.
  * tests/e2e/conftest.py: .env 자동 로드 + provider별 skip 조건 (Ollama는
    LLM_API_KEY 불필요).
  * tests/e2e/test_phase0_full_pipeline.py: provider별 키 분기,
    HDBSCAN 클러스터링이 동작하도록 fixture 페이로드 16문장으로 확장.

- vendored OntoCast 버그 수정 3건 (VENDORED_MODIFICATIONS.md 기록):
  * agent/render_ontology.py: render_ontology_fresh()의 .format() 호출에 누락된
    ontology_prefix 인자 추가 (Bootstrap 단계에서 KeyError: 'ontology_prefix').
  * stategraph/node_factories.py: render_ontology/render_facts 노드의
    state.model_copy(deep=True)로 budget_tracker가 deep-copy되어 root state의
    BudgetTracker가 영원히 0인 채로 남던 버그 수정. 원본 인스턴스 공유로 변경.

- 문서 갱신
  README.md (Phase 0.7 부분완료 + ont_platform 폴더 이름),
  docs/phases/PHASE0_ACCEPTANCE_GATE.md (검증 이력 + Ollama/LM Studio 옵션),
  .env.example (LM Studio/Ollama/OpenAI 세 옵션 명시).

검증
- unit + integration 26/26 통과.
- e2e (LM Studio + Qwen3-8B / DeepSeek-R1-Distill-Qwen-7B): 워크플로우 끝까지
  실행 + 5번 LLM 호출 + LangGraph 전 노드 traceable 확인. 7-8B 로컬 모델은
  strict structured output(Turtle RDF in JSON) 한계로 ontology/facts TTL 자동
  생성 부분 성공. 클라우드 LLM 환경에서 재검증 필요.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-05-14 09:05:24 +09:00

205 lines
7.2 KiB
Python

"""Regression tests for the OntoCast `convert_document` agent.
Covers the multi-file corpus extension described in
`docs/통합설계서.md` §5 Phase 0 and OntoCast 분석 §13.1 / §21.1.
Original behavior: the loop overwrote `state.input_text` on every iteration,
so multi-file input silently lost all but the last file. The fix accumulates
into one corpus with explicit file-boundary separators while keeping
single-file behavior byte-identical.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
from types import SimpleNamespace
import pytest
REPO_ROOT = Path(__file__).resolve().parents[2]
VENDORED_ONTOCAST = REPO_ROOT / "vendored" / "ontocast"
if str(VENDORED_ONTOCAST) not in sys.path:
sys.path.insert(0, str(VENDORED_ONTOCAST))
import sys
import ontocast.agent # __init__.py 실행으로 서브모듈이 sys.modules에 등록됨 # noqa: E402
convert_document_module = sys.modules["ontocast.agent.convert_document"]
from ontocast.onto.enum import Status # noqa: E402
class _StubConverter:
"""Minimal stand-in for `ConverterTool`.
Accepts a fake PDF/DOCX file (any bytes) and returns a dict shaped like
the real converter's output.
"""
supported_extensions = {".pdf", ".docx"}
def __init__(self, mapping: dict[bytes, str]) -> None:
self._mapping = mapping
def __call__(self, file_content: bytes) -> dict[str, str]:
return {"text": self._mapping[file_content]}
def _make_state(files: dict[str, bytes]) -> SimpleNamespace:
"""Build a lightweight stand-in for `AgentState`.
We intentionally avoid constructing the real Pydantic model here — its
initialization touches many unrelated fields and tools. We only mirror
the attributes that `convert_document` reads or writes.
"""
captured_text: list[str] = []
def set_text(text: str) -> None:
captured_text.append(text)
state = SimpleNamespace(
files=files,
status=None,
input_text="",
ontology_user_instruction="",
facts_user_instruction="",
source_url=None,
set_text=set_text,
_captured_text=captured_text,
)
return state
def _make_tools(converter_mapping: dict[bytes, str]) -> SimpleNamespace:
return SimpleNamespace(converter=_StubConverter(converter_mapping))
# ─── Case 1: single PDF — output must equal the file's text verbatim ─────
def test_single_pdf_passes_through_unchanged() -> None:
pdf_bytes = b"%PDF-1.4 fake"
state = _make_state({"sample.pdf": pdf_bytes})
tools = _make_tools({pdf_bytes: "Hello from PDF."})
result = convert_document_module.convert_document(state, tools)
assert result.status == Status.SUCCESS
assert state._captured_text == ["Hello from PDF."]
# ─── Case 2: single JSON — text + corpus metadata flows through ──────────
def test_single_json_extracts_metadata_and_text() -> None:
payload = {
"text": "Body of the article.",
"url": "https://example.com/a",
"ontology_user_instruction": "Focus on organizations.",
"facts_user_instruction": "Extract person-org links.",
}
state = _make_state({"a.json": json.dumps(payload).encode("utf-8")})
tools = _make_tools({})
result = convert_document_module.convert_document(state, tools)
assert result.status == Status.SUCCESS
assert state._captured_text == ["Body of the article."]
assert state.source_url == "https://example.com/a"
assert state.ontology_user_instruction == "Focus on organizations."
assert state.facts_user_instruction == "Extract person-org links."
# ─── Case 3: multiple PDFs — both bodies survive with boundary marker ────
def test_multiple_pdfs_are_concatenated_with_boundary() -> None:
"""The legacy bug: only the last file's text survived. After the fix,
both texts appear in the corpus separated by `=== File: <name> ===`."""
pdf_a = b"%PDF-1.4 A"
pdf_b = b"%PDF-1.4 B"
state = _make_state({"a.pdf": pdf_a, "b.pdf": pdf_b})
tools = _make_tools({pdf_a: "Text A.", pdf_b: "Text B."})
convert_document_module.convert_document(state, tools)
corpus = state._captured_text[-1]
assert "Text A." in corpus
assert "Text B." in corpus
assert "=== File: a.pdf ===" in corpus
assert "=== File: b.pdf ===" in corpus
# Ordering: a.pdf before b.pdf (insertion order preserved)
assert corpus.index("Text A.") < corpus.index("Text B.")
# ─── Case 4: multiple JSONs — first-wins for corpus metadata ─────────────
def test_multiple_jsons_keep_first_metadata() -> None:
"""`ontology_user_instruction`, `facts_user_instruction`, `source_url`
are corpus-level singletons. The first JSON file that provides each
wins; later JSONs do not overwrite."""
payload_a = {
"text": "Body A.",
"url": "https://first.example/a",
"ontology_user_instruction": "First instruction.",
"facts_user_instruction": "First facts.",
}
payload_b = {
"text": "Body B.",
"url": "https://second.example/b",
"ontology_user_instruction": "Second instruction (must be ignored).",
"facts_user_instruction": "Second facts (must be ignored).",
}
state = _make_state(
{
"a.json": json.dumps(payload_a).encode("utf-8"),
"b.json": json.dumps(payload_b).encode("utf-8"),
}
)
tools = _make_tools({})
convert_document_module.convert_document(state, tools)
assert state.source_url == "https://first.example/a"
assert state.ontology_user_instruction == "First instruction."
assert state.facts_user_instruction == "First facts."
corpus = state._captured_text[-1]
assert "Body A." in corpus and "Body B." in corpus
# ─── Case 5: unsupported extension fails fast ────────────────────────────
def test_unsupported_extension_returns_failed() -> None:
state = _make_state({"weird.xyz": b"???"})
tools = _make_tools({})
result = convert_document_module.convert_document(state, tools)
assert result.status == Status.FAILED
# ─── Case 6: empty files dict is a no-op success ─────────────────────────
def test_empty_files_is_noop_success() -> None:
state = _make_state({})
tools = _make_tools({})
result = convert_document_module.convert_document(state, tools)
assert result.status == Status.SUCCESS
assert state._captured_text == [] # set_text never called
# ─── Case 7: mixed PDF + JSON in one corpus ──────────────────────────────
def test_mixed_pdf_and_json_combine_with_boundaries() -> None:
pdf_bytes = b"%PDF-1.4 mix"
json_payload = {"text": "JSON body."}
state = _make_state(
{
"a.pdf": pdf_bytes,
"b.json": json.dumps(json_payload).encode("utf-8"),
}
)
tools = _make_tools({pdf_bytes: "PDF body."})
convert_document_module.convert_document(state, tools)
corpus = state._captured_text[-1]
assert "=== File: a.pdf ===" in corpus
assert "=== File: b.json ===" in corpus
assert "PDF body." in corpus
assert "JSON body." in corpus