Files
AI/참고/ontology_crawler_open_source_integration_plan.md

655 lines
7.9 KiB
Markdown
Raw Normal View History

2026-05-12 19:40:31 +09:00
# Ontology Crawler Platform 개선 전략
현재 프로젝트는 폐기 대상이 아니라, 구조는 유지하면서 핵심 파이프라인을 교체 및 강화해야 하는 단계이다.
현재 가장 큰 문제는:
- 웹페이지 본문 추출 실패
- footer/nav/배송정보/이벤트 오염
- LLM JSON 출력 실패
- ontology validation 부재
- fallback 규칙 오염
- entity/page classification 부재
이며,
이 문제는 단순히 LLM 모델 교체로 해결되지 않는다.
핵심은:
```text
좋은 본문을 추출하고
→ 안정적으로 구조화하고
→ 검증된 ontology claim만 저장하는 것
```
이다.
---
# 현재 다운로드한 OSS 역할 정리
| 프로젝트 | 역할 | 현재 프로젝트에서 사용 목적 |
|---|---|---|
| playwright | 브라우저 렌더링 | JS 렌더링 / SPA / 실제 페이지 획득 |
| crawl4ai | AI 친화 크롤링 | markdown 변환 / content cleaning |
| trafilatura | 본문 추출 | footer/nav 제거 / main content extraction |
| instructor | LLM JSON 안정화 | strict schema output |
| guardrails | LLM validation | garbage output reject |
| ontocast | ontology 구조 참고 | entity/relation/triple 구조 |
| neo4j-graphrag | graph 구조 참고 | ontology graph 저장 |
| knowledge_agent | agent loop 참고 | recursive crawl / relevance |
| OpenDeepResearcher | research flow 참고 | 탐색 loop 설계 |
| firecrawl | AI friendly crawling 참고 | markdown cleaning 구조 |
---
# 프로젝트 방향 재정의
현재 프로젝트는 단순 crawler가 아니라:
```text
Semantic Ontology Research Platform
```
방향으로 정의한다.
최종 목표:
```text
→ 의미 추출
→ entity/relation 생성
→ ontology mapping
→ graph knowledge 구축
→ semantic exploration
```
---
# 현재 문제 분석
# 문제 1. 페이지 전체를 분석하고 있음
현재:
- footer
- 배송 국가 목록
- 이벤트 배너
- 공지
- navigation
- 게시판 링크
- 정렬 텍스트
까지 ontology claim 생성에 포함되고 있다.
결과:
```text
CAFE24 = 브랜드
Green = accord
낮은가격 = entity
상품수 = entity
```
같은 잘못된 결과 발생.
이 문제는:
```text
본문 추출 실패
```
문제이다.
---
# 문제 2. AI extractor가 실제로 거의 실패 중
현재 로그:
```text
Expecting value: line 1 column 1
AI returned no usable entities or claims
```
의미:
- JSON 반환 실패
- 빈 문자열
- markdown/codeblock
- 설명문 반환
- malformed JSON
등이 발생 중.
현재는 fallback rule이 실제 ontology claim처럼 저장되고 있음.
이것은 매우 위험하다.
---
# 문제 3. Entity type 분리 없음
현재:
- Product
- Event
- Promotion
- Notice
- BrandStory
- CommunityPost
가 모두 같은 레벨에서 처리된다.
그래서 ontology graph가 오염된다.
---
# 문제 4. Validation 없음
현재 저장되는 값:
```text
value
accord
keyword
""
```
이것은 ontology 데이터가 아니다.
Placeholder reject 필요.
---
# 개선 방향 전체 구조
현재 구조를 아래 파이프라인으로 재구성한다.
```text
[Crawler]
→ [DOM/MainContent Extractor]
→ [Page Classifier]
→ [Structured Page Model]
→ [LLM Extraction]
→ [Schema Validation]
→ [Claim Validation]
→ [Ontology Mapping]
→ [Candidate Review]
→ [Graph Merge]
```
---
# 1단계 — Playwright 도입
사용 OSS:
- playwright
목표:
```text
실제 브라우저 기반 페이지 획득
```
도입 목적:
- JS 렌더링
- lazy loading 대응
- dynamic page 대응
- 실제 DOM 안정화
- SPA 대응
구현:
```text
1. 페이지 이동
2. network idle 대기
3. 특정 selector 대기
4. HTML snapshot 저장
5. canonical URL 저장
6. error/captcha 감지
```
추가:
- blocked page 감지
- unavailable page 감지
- region redirect 감지
예:
```text
Page unavailable
Access denied
captcha
Reference ID
```
등 발견 시:
```text
crawl_failed
```
상태 저장.
---
# 2단계 — Trafilatura 본문 추출
사용 OSS:
- trafilatura
현재 프로젝트에서 가장 중요한 단계.
목표:
```text
웹페이지에서 실제 읽을 본문만 추출
```
제거 대상:
- footer
- nav
- shipping info
- recommendation
- banner
- copyright
- category
- board links
- country lists
출력:
```text
clean_text
clean_html
main_content
```
현재 문제:
```text
배송 국가 목록이 accord로 추출됨
```
같은 문제를 여기서 해결.
---
# 3단계 — Crawl4AI 도입
사용 OSS:
- crawl4ai
목적:
```text
LLM 친화 markdown 생성
```
흐름:
```text
HTML
→ markdown
→ cleaned markdown
→ semantic blocks
```
기대 효과:
- 긴 HTML 감소
- token 감소
- semantic extraction 안정화
- hallucination 감소
---
# 4단계 — Page Classification 추가
현재 가장 부족한 구조.
도입:
```text
PageType classifier
```
분류:
- ProductPage
- BrandStoryPage
- PromotionPage
- CommunityPage
- EventPage
- CategoryPage
- NoticePage
- UnknownPage
예:
```text
/product/detail
→ ProductPage
/board/free
→ CommunityPage
/shopinfo
→ BrandStoryPage
```
중요:
ProductPage가 아니면:
```text
price extraction 금지
```
등 claim 제한 필요.
---
# 5단계 — Instructor 도입
사용 OSS:
- instructor
목표:
```text
LLM output strict JSON 강제
```
현재 문제:
```text
Expecting value
```
해결 목적.
예:
```python
class ProductClaim(BaseModel):
product_name: str
brand: str
price: int
top_notes: list[str]
```
LLM이 schema에 맞지 않으면 reject.
---
# 6단계 — Guardrails Validation
사용 OSS:
- guardrails
목표:
```text
garbage ontology reject
```
Reject 대상:
```text
value
accord
keyword
""
null
unknown
n/a
```
추가:
- too short reject
- meaningless token reject
- duplicated relation reject
- invalid entity reject
---
# 7단계 — Claim 상태 분리
현재:
```text
rule 결과가 ontology claim으로 저장됨
```
문제.
변경:
```text
candidate_claim
validated_claim
rejected_claim
ai_claim
rule_candidate
```
분리.
rule 결과는:
```text
candidate 상태
```
로만 저장.
---
# 8단계 — Ontology Graph 구조 개선
참고 OSS:
- ontocast
- neo4j-graphrag
목표:
```text
Entity
→ Relation
→ Triple
→ Graph
```
예:
```text
[Product] Cotton Hug
├─ hasBrand → Forment
├─ hasPrice → 49000 KRW
├─ hasTopNote → Pink Pepper
├─ suitableForSeason → Summer
```
현재처럼:
```text
CAFE24
낮은가격
상품수
```
등은 graph merge 금지.
---
# 9단계 — Recursive Research Agent
후반 단계.
참고 OSS:
- knowledge_agent
- OpenDeepResearcher
목표:
```text
탐색
→ relevance 판단
→ 링크 확장
→ ontology 성장
```
예:
```text
브랜드 페이지 발견
→ 제품 페이지 탐색
→ 리뷰 페이지 탐색
→ 노트 정보 탐색
→ 경쟁 브랜드 탐색
```
---
# 추천 구현 순서
# Phase 1
핵심 안정화.
```text
Playwright
Trafilatura
Page Cleaner
```
우선.
현재 프로젝트의 가장 중요한 문제는:
```text
LLM 이전 단계 실패
```
이기 때문.
---
# Phase 2
```text
Instructor
Guardrails
```
도입.
LLM JSON 안정화.
---
# Phase 3
```text
Page Classification
Claim Validation
Candidate Separation
```
추가.
---
# Phase 4
```text
Neo4j Graph Merge
Ontology Relation
Semantic Query
```
추가.
---
# Phase 5
```text
Knowledge Agent
Recursive Research
Semantic Expansion
```
추가.
---
# 최종 방향
이 프로젝트는:
```text
단순 크롤러
```
가 아니라:
```text
Semantic Knowledge Construction Platform
```
으로 발전해야 한다.
핵심은:
```text
웹을 읽는 것
```
이 아니라:
```text
웹의 의미 구조를 구축하는 것
```
이다.
따라서:
- 좋은 본문 추출
- strict validation
- ontology integrity
- semantic relation
- graph quality
가 가장 중요하다.
현재 문제는:
```text
LLM 성능 부족
```
보다 먼저:
```text
좋은 데이터를 제대로 추출하지 못하는 것
```
이다.