Files
AI/ontology_platform/docs/semantic_page_classification_codex_spec.md

2266 lines
38 KiB
Markdown
Raw Permalink Normal View History

2026-05-22 20:16:28 +09:00
# Semantic Page Classification Layer 설계 및 구현 지시서
## 0. 문서 목적
본 문서는 온톨로지 플랫폼의 `page_classifier.py`를 단순 URL/텍스트 기반 페이지 분류기에서 **범용 Semantic Page Understanding Layer**로 확장하기 위한 작업 지시서이다.
현재 시스템은 다음과 같은 제한적인 page type만 다룬다.
```txt
ProductPage
CategoryPage
SearchPage
BoardPage
NoticePage
BrandStoryPage
PromotionPage
UnknownPage
```
이 구조는 쇼핑몰 일부 페이지를 분류하는 데는 사용할 수 있으나, 범용 온톨로지 구축 플랫폼에는 부족하다.
범용 플랫폼은 인터넷에 존재하는 다양한 페이지를 다음 관점으로 분류할 수 있어야 한다.
```txt
Page Domain
Page Archetype
Semantic Page Type
Main Entity Type
Action Intent
Graph Role
Analyze Strategy
LLM Policy
Confidence
Evidence
```
즉, 목표는 단순히 `ProductPage` 같은 enum 하나를 맞히는 것이 아니라, 페이지의 의미적 역할을 여러 축으로 해석하고, 온톨로지 그래프 구축에 필요한 분석 전략까지 결정하는 것이다.
---
## 1. 현재 구조의 문제점
### 1.1 Page Type이 너무 적다
현재 분류 항목은 대부분 commerce 또는 게시판 중심이다.
```txt
ProductPage
CategoryPage
SearchPage
BoardPage
BrandStoryPage
```
하지만 인터넷에는 다음과 같은 페이지가 존재한다.
```txt
ArticlePage
NewsPage
BlogPostPage
FAQPage
QAPage
WikiPage
DocumentationPage
JobPostingPage
CoursePage
VideoPage
ProfilePage
LocalBusinessPage
RealEstateListingPage
PricingPage
CheckoutPage
LoginPage
TermsPage
PrivacyPolicyPage
APIReferencePage
DatasetPage
ResearchPaperPage
```
현 구조에서는 이들을 대부분 `UnknownPage` 또는 부정확한 기존 타입으로 분류하게 된다.
---
### 1.2 URL 휴리스틱 의존도가 높다
현재 방식은 대체로 다음과 같은 구조다.
```python
if "/product/" in url:
return "ProductPage"
if "/board/" in url:
return "BoardPage"
```
이 방식은 다음 환경에서 쉽게 깨진다.
```txt
SPA
Headless Commerce
Dynamic Route
Query 기반 페이지
다국어 URL
짧은 URL
해시 라우팅
CMS 기반 자동 생성 페이지
AI 생성 페이지
```
예:
```txt
/p/12345
/x/abc
/node/987
/view?id=123
/ko/contents/123
```
이 URL만 보고는 페이지 의미를 알 수 없다.
---
### 1.3 CategoryPage, SearchPage, BoardPage를 너무 쉽게 skip한다
현재 구조는 보통 다음 흐름이다.
```txt
classify_page()
-> should_analyze_page()
-> analyze_page_types에 없으면 Extractor 실행 안 함
```
이 때문에 `CategoryPage`, `SearchPage`, `BoardPage`는 기본적으로 분석 대상에서 빠질 가능성이 높다.
그러나 범용 온톨로지 플랫폼에서는 이 페이지들이 중요하다.
```txt
CategoryPage -> 카테고리 계층, 상품 목록, taxonomy 관계
SearchPage -> 검색 의도, 결과 후보, query-result 관계
BoardPage -> 게시글 목록, 질문/답변/토론 구조
ListingPage -> entity collection, relation hub
ArchivePage -> 시간 축 기반 콘텐츠 구조
TagPage -> topic-entity 관계
```
따라서 “분석 여부”는 단순 boolean이 아니라, 페이지 타입별 분석 전략으로 분리해야 한다.
---
### 1.4 UnknownPage를 버리면 안 된다
현재는 분류 실패 시 `UnknownPage`로 두고 사실상 분석에서 제외될 가능성이 크다.
하지만 범용 플랫폼에서 `UnknownPage`는 새로운 페이지 패턴을 발견하는 출발점이다.
예:
```txt
UnknownPage Cluster A:
- 향수 노트 비교표가 있음
- 여러 상품을 향 계열별로 비교
- 일반 ProductPage도 CategoryPage도 아님
새 후보:
PerfumeNoteComparisonPage
```
따라서 Unknown은 폐기 대상이 아니라, clustering 및 taxonomy 확장 후보로 저장해야 한다.
---
## 2. 목표 아키텍처
기존 구조:
```txt
URL / title / text / html
-> page_classifier.py
-> page_type
-> should_analyze_page()
```
개선 구조:
```txt
Raw Page Snapshot
-> Signal Extraction
-> Evidence Scoring
-> Domain Classification
-> Archetype Classification
-> Entity Type Classification
-> Action Intent Classification
-> Graph Role Assignment
-> Semantic Page Type Decision
-> Analyze Strategy Decision
-> LLM Policy Decision
-> Unknown Pattern Storage / Clustering
```
---
## 3. 핵심 개념
### 3.1 Page Domain
페이지가 속한 큰 의미 영역이다.
```txt
Commerce
Editorial
Community
Knowledge
Corporate
Local
Education
Jobs
Media
Software
Finance
Government
Healthcare
Transaction
System
Unknown
```
---
### 3.2 Page Archetype
페이지의 구조적 역할이다.
```txt
Home
Landing
Detail
Listing
Collection
SearchResult
Profile
Article
Thread
Form
Transaction
Dashboard
Document
Media
Error
SystemResource
Unknown
```
---
### 3.3 Semantic Page Type
구체적인 페이지 타입이다.
예:
```txt
ProductDetailPage
CategoryListingPage
ArticlePage
ForumThreadPage
FAQPage
JobPostingPage
LocalBusinessPage
CheckoutPage
LoginPage
DocumentationPage
```
---
### 3.4 Main Entity Type
페이지가 중심으로 삼는 엔티티이다.
```txt
Product
Service
Article
NewsArticle
Person
Organization
Place
Event
JobPosting
Course
Question
Answer
Review
Dataset
SoftwareApplication
MediaObject
Recipe
RealEstateProperty
MedicalCondition
LegalDocument
FinancialProduct
UnknownEntity
```
---
### 3.5 Action Intent
페이지가 사용자를 유도하는 행동이다.
```txt
Read
Buy
Subscribe
Reserve
Book
Apply
Download
Watch
Listen
Search
Compare
Filter
Ask
Answer
Comment
Review
Login
Register
Pay
Contact
Navigate
Learn
Verify
Configure
Manage
```
---
### 3.6 Graph Role
온톨로지 그래프 안에서 이 페이지가 수행하는 역할이다.
```txt
EntityAnchor
RelationHub
NavigationHub
CollectionHub
SearchHub
TransactionOnly
PolicySource
ClaimSource
ProfileAnchor
MediaAnchor
ReferenceSource
SystemResource
NoisePage
UnknownPattern
```
---
### 3.7 Analyze Strategy
페이지를 어떻게 분석할지 결정하는 전략이다.
```txt
AnalyzeFull
AnalyzeStructureOnly
AnalyzeEntityOnly
AnalyzeRelationsOnly
AnalyzeMetadataOnly
AnalyzeDocumentOnly
AnalyzeDiscoveryOnly
SkipProtected
SkipNoise
```
---
### 3.8 LLM Policy
LLM 사용 여부 및 사용 범위다.
```txt
LLMFull
LLMLight
LLMForAmbiguityOnly
RuleOnly
NoLLM
Skip
```
---
## 4. Page Taxonomy v1
아래 taxonomy는 초기 버전이다. 구현 시 enum 또는 문자열 상수로 관리한다.
---
## 4.1 Site / Navigation 계열
```txt
HomePage
LandingPage
PortalPage
SectionHomePage
CategoryPage
SubcategoryPage
TagPage
TopicPage
CollectionPage
ArchivePage
SitemapPage
DirectoryPage
IndexPage
SearchPage
SearchResultsPage
FilteredResultsPage
PaginationPage
LocaleSelectorPage
LanguageRedirectPage
RedirectPage
NotFoundPage
ErrorPage
MaintenancePage
ComingSoonPage
RobotsBlockedPage
```
### 주요 신호
```txt
many internal links
breadcrumb
category tree
pagination
tag links
archive dates
search form
sitemap XML or sitemap-like links
locale links
error status text
```
### 온톨로지 역할
```txt
사이트 구조 파악
카테고리 계층 파악
내부 링크 그래프 구축
탐색 우선순위 결정
```
---
## 4.2 Commerce / Marketplace 계열
```txt
ProductDetailPage
ProductVariantPage
ProductBundlePage
ProductComparisonPage
ProductReviewPage
ProductQnAPage
ProductManualPage
ProductSpecPage
CategoryListingPage
ProductListingPage
BrandCatalogPage
SellerStorePage
MarketplaceListingPage
SearchProductResultsPage
DealPage
SalePage
CouponPage
PromotionPage
CampaignLandingPage
SubscriptionPlanPage
PricingPage
CartPage
CheckoutPage
PaymentPage
OrderPage
OrderConfirmationPage
OrderTrackingPage
WishlistPage
GiftCardPage
StoreLocatorPage
InventoryAvailabilityPage
AuctionPage
RentalProductPage
BookingProductPage
ServiceProductPage
```
### 주요 신호
```txt
schema.org Product
schema.org Offer
schema.org AggregateRating
price
currency
availability
add to cart
buy now
variant selector
quantity selector
SKU
brand
product image gallery
reviews
rating
shipping
return policy
repeated product cards
filters
sort control
pagination
```
### 온톨로지 역할
```txt
Product 엔티티 생성
Brand 관계 생성
Category 관계 생성
Offer / Price / Availability 추출
Review / Rating 관계 추출
RelatedProduct 관계 추출
```
---
## 4.3 Editorial / Article / Publishing 계열
```txt
ArticlePage
NewsArticlePage
BlogPostPage
OpinionPage
EditorialPage
InterviewPage
ReportPage
ColumnPage
PressArticlePage
MagazinePage
GuidePage
TutorialPage
HowToPage
RecipePage
CaseStudyPage
WhitePaperPage
ResearchSummaryPage
StoryPage
ChapterPage
SeriesPage
AuthorArticleListPage
PaywalledArticlePage
SponsoredContentPage
```
### 주요 신호
```txt
schema.org Article
schema.org NewsArticle
schema.org BlogPosting
headline
author
publisher
datePublished
dateModified
articleBody
byline
section
tags
hero image
related articles
paywall marker
```
### 온톨로지 역할
```txt
Article / Topic / Author / Publisher 엔티티 생성
about 관계 생성
citation / source 관계 생성
temporal coverage 추출
claim 후보 추출
```
---
## 4.4 Community / UGC 계열
```txt
ForumHomePage
ForumBoardPage
ForumThreadPage
DiscussionPage
CommentThreadPage
QAPage
FAQPage
ReviewPage
UserReviewPage
CommunityPostPage
SocialPostPage
TimelinePage
FeedPage
UserProfilePage
CreatorProfilePage
GroupPage
CommunityPage
PollPage
PetitionPage
RankingPage
LeaderboardPage
ReputationPage
BadgePage
```
### 주요 신호
```txt
question
answer
accepted answer
comments
reply
thread
votes
likes
author profile
user avatar
posted date
edited date
review rating
FAQ accordion
Q&A structured data
```
### 온톨로지 역할
```txt
Question / Answer 엔티티 생성
User / Author 관계 생성
Thread 관계 생성
Claim / Opinion 분리
Reputation / Vote / AcceptedAnswer 추출
```
---
## 4.5 Knowledge / Reference / Documentation 계열
```txt
WikiPage
EncyclopediaPage
GlossaryPage
DefinitionPage
ReferencePage
DocumentationPage
DeveloperDocsPage
APIDocumentationPage
APIReferencePage
SDKDocumentationPage
ManualPage
SpecificationPage
StandardPage
ProtocolPage
ChangelogPage
ReleaseNotesPage
ErrorCodePage
TroubleshootingPage
KnowledgeBaseArticlePage
DatasetPage
DataCatalogPage
ResearchPaperPage
PatentPage
CitationPage
BibliographyPage
```
### 주요 신호
```txt
definition
table of contents
code block
API method
parameter table
version
endpoint
changelog
release notes
specification
standard
citation
references
dataset metadata
```
### 온톨로지 역할
```txt
Concept / Term / Definition 추출
API / Method / Parameter 관계 추출
Version 관계 추출
Dataset metadata 추출
Reference graph 생성
```
---
## 4.6 Corporate / Organization 계열
```txt
AboutPage
CompanyPage
BrandStoryPage
MissionPage
VisionPage
HistoryPage
TeamPage
FounderPage
LeadershipPage
ContactPage
LocationPage
BranchPage
InvestorRelationsPage
IRPage
FinancialReportPage
PressReleasePage
MediaKitPage
PartnershipPage
FranchisePage
CareersHomePage
JobPostingPage
RecruitPage
CulturePage
LegalPage
TermsPage
PrivacyPolicyPage
CookiePolicyPage
AccessibilityPage
CompliancePage
SecurityPage
TrustCenterPage
```
### 주요 신호
```txt
about us
company
mission
vision
history
team
founder
leadership
contact
address
investor relations
press release
careers
privacy policy
terms of service
cookie policy
security
compliance
```
### 온톨로지 역할
```txt
Organization 엔티티 생성
Founder / Location / Contact 관계 생성
Policy 문서 분류
법적/계약적 문장 추출
채용 정보 추출
```
---
## 4.7 Local / Place / Travel / Real Estate 계열
```txt
PlaceDetailPage
LocalBusinessPage
RestaurantPage
MenuPage
HotelPage
RoomPage
VacationRentalPage
TravelDestinationPage
AttractionPage
ItineraryPage
MapPage
MapSearchResultsPage
RealEstateListingPage
PropertyDetailPage
PropertySearchResultsPage
AgentProfilePage
OpenHousePage
ReservationPage
BookingPage
AvailabilityCalendarPage
TransportRoutePage
FlightPage
TrainPage
BusRoutePage
```
### 주요 신호
```txt
address
geo coordinates
map
opening hours
menu
reservation
booking
room availability
travel dates
property price
bedrooms
bathrooms
area
agent
route
schedule
```
### 온톨로지 역할
```txt
Place 엔티티 생성
Address / Geo / OpeningHours 추출
Reservation 가능성 판단
Nearby 관계 생성
Availability 추출
```
---
## 4.8 Education / Learning 계열
```txt
CourseDetailPage
CourseListPage
CurriculumPage
LessonPage
LecturePage
TutorialPage
AssignmentPage
QuizPage
ExamPage
FlashcardPage
EducationQAPage
MathSolverPage
SchoolPage
UniversityPage
ProgramPage
DegreePage
CertificationPage
InstructorProfilePage
LearningPathPage
```
### 주요 신호
```txt
course
lesson
curriculum
instructor
learning objective
assignment
quiz
exam
certificate
degree
program
tuition
syllabus
```
### 온톨로지 역할
```txt
Course / Lesson / Instructor 엔티티 생성
Prerequisite 관계 생성
LearningObjective 추출
Question / Answer / Solution 구조화
```
---
## 4.9 Jobs / Career 계열
```txt
JobPostingPage
JobSearchResultsPage
CompanyJobsPage
CareerCategoryPage
ApplicationFormPage
RecruitmentLandingPage
EmployerProfilePage
EmployerReviewPage
SalaryPage
InterviewReviewPage
BenefitsPage
InternshipPage
FreelanceGigPage
```
### 주요 신호
```txt
job title
employment type
salary
location
remote
apply
requirements
responsibilities
benefits
company
recruiter
deadline
```
### 온톨로지 역할
```txt
Job 엔티티 생성
Employer 관계 생성
Location / Salary / EmploymentType 추출
Skill requirement 추출
```
---
## 4.10 Media / Entertainment 계열
```txt
VideoPage
VideoWatchPage
LiveStreamPage
PodcastPage
EpisodePage
MusicTrackPage
AlbumPage
ArtistPage
MoviePage
TVSeriesPage
TVEpisodePage
GameDetailPage
GameGuidePage
ImagePage
ImageGalleryPage
PhotoStoryPage
MediaGalleryPage
DownloadMediaPage
StreamingChannelPage
```
### 주요 신호
```txt
video player
audio player
duration
episode
season
album
artist
track
movie
trailer
live
stream
gallery
image grid
download
```
### 온톨로지 역할
```txt
MediaObject 엔티티 생성
Creator / Performer / Publisher 관계 생성
Duration / Episode / Series 관계 추출
License / UsageInfo 추출
```
---
## 4.11 Software / SaaS / App 계열
```txt
SoftwareProductPage
SaaSProductPage
FeaturePage
PricingPage
IntegrationPage
PluginPage
ExtensionPage
AppStoreListingPage
PackagePage
RepositoryPage
ReleasePage
ChangelogPage
IssuePage
PullRequestPage
DocumentationPage
APIReferencePage
StatusPage
DashboardPage
SettingsPage
AdminPage
LoginPage
SignupPage
OnboardingPage
BillingPage
UsageReportPage
```
### 주요 신호
```txt
software
app
SaaS
feature
pricing
integration
plugin
extension
repository
release
changelog
issue
pull request
status
dashboard
settings
billing
API
SDK
```
### 온톨로지 역할
```txt
Software / Version / Feature 엔티티 생성
Dependency 관계 생성
Release 관계 생성
Issue / PR / Commit 관계 추출
```
---
## 4.12 Finance / Legal / Government 계열
```txt
BankProductPage
LoanPage
CreditCardPage
InsuranceProductPage
InvestmentProductPage
StockQuotePage
CryptoAssetPage
FinancialReportPage
TaxInfoPage
GovernmentServicePage
PublicNoticePage
RegulationPage
LawPage
CourtCasePage
LegalArticlePage
PolicyPage
FormPage
ApplicationPage
PermitPage
LicensePage
PublicDataPage
ProcurementPage
TenderPage
```
### 주요 신호
```txt
interest rate
APR
loan
credit card
insurance
investment
stock quote
financial statement
tax
government
regulation
law
court
policy
permit
license
tender
procurement
```
### 온톨로지 역할
```txt
Regulation / Policy / Law 엔티티 생성
Obligation / Prohibition / Permission 추출
Institution 관계 생성
Form requirement 추출
```
---
## 4.13 Healthcare / Medical 계열
```txt
MedicalArticlePage
ConditionPage
SymptomPage
TreatmentPage
DrugPage
SupplementPage
DoctorProfilePage
HospitalPage
ClinicPage
AppointmentPage
InsuranceCoveragePage
ClinicalTrialPage
MedicalFAQPage
HealthCalculatorPage
EmergencyInfoPage
```
### 주요 신호
```txt
condition
symptom
treatment
drug
dosage
side effect
doctor
hospital
clinic
appointment
clinical trial
insurance coverage
emergency
```
### 온톨로지 역할
```txt
Condition / Treatment / Drug 엔티티 생성
Symptom 관계 생성
Medical claim 추출
Source reliability 분리
```
---
## 4.14 Transaction / Account / Protected 계열
```txt
LoginPage
SignupPage
PasswordResetPage
AccountPage
ProfileSettingsPage
NotificationPage
MessageInboxPage
CartPage
CheckoutPage
PaymentPage
SubscriptionManagementPage
BillingPage
InvoicePage
OrderHistoryPage
UploadPage
DownloadPage
FormPage
SurveyPage
ConsentPage
AgeGatePage
CaptchaPage
PaywallPage
AccessDeniedPage
SessionExpiredPage
```
### 주요 신호
```txt
login form
password field
signup
reset password
account settings
payment fields
checkout
billing
invoice
consent
captcha
age gate
access denied
session expired
```
### 온톨로지 역할
```txt
대부분 분석 제외
개인정보 보호
크롤링 중단 또는 제한
거래 흐름만 메타 수준으로 기록
```
---
## 4.15 System / Technical / Machine-readable 계열
```txt
RSSFeedPage
AtomFeedPage
XMLSitemapPage
RobotsTxtPage
ManifestPage
OpenSearchDescriptionPage
JSONEndpointPage
APIEndpointPage
GraphQLEndpointPage
WebhookEndpointPage
FileDownloadPage
PDFDocumentPage
CSVDocumentPage
XMLDocumentPage
ImageAssetPage
VideoAssetPage
FontAssetPage
ScriptAssetPage
StylesheetAssetPage
```
### 주요 신호
```txt
content-type
xml
json
rss
atom
sitemap
robots.txt
manifest
API response
file extension
download headers
```
### 온톨로지 역할
```txt
크롤링 정책 파악
사이트 구조 파악
데이터 소스 발견
문서형 리소스 별도 파서로 전달
```
---
## 5. Signal Extraction 설계
`page_classifier.py`가 직접 모든 것을 처리하지 말고, signal extractor를 분리한다.
권장 파일 구조:
```txt
ontology_platform/
classifier/
page_classifier.py
page_type_taxonomy.py
page_signals.py
page_signal_extractor.py
page_type_scorer.py
page_analysis_policy.py
adaptive_page_classifier.py
```
기존 프로젝트 구조에 맞춰 경로는 조정해도 된다.
---
## 5.1 Raw Page Snapshot
분류 함수 입력은 다음 정보를 받을 수 있어야 한다.
```python
@dataclass
class RawPageSnapshot:
url: str
final_url: str | None
status_code: int | None
content_type: str | None
title: str | None
text: str | None
html: str | None
rendered_html: str | None
metadata: dict
open_graph: dict
twitter_card: dict
json_ld: list[dict]
microdata: list[dict]
rdfa: list[dict]
headings: list[str]
links: list[dict]
images: list[dict]
forms: list[dict]
buttons: list[str]
inputs: list[dict]
tables: list[dict]
breadcrumbs: list[str]
source_zones: list[str]
screenshot_path: str | None
```
초기 구현에서는 모든 필드가 없어도 된다.
없는 값은 `None` 또는 빈 리스트로 처리한다.
---
## 5.2 Page Signals
추출 결과는 다음 형태로 관리한다.
```python
@dataclass
class PageSignals:
# structured data
schema_types: set[str]
og_type: str | None
twitter_card_type: str | None
# commerce
has_price: bool
has_currency: bool
has_cart_button: bool
has_buy_button: bool
has_variant_selector: bool
has_sku: bool
has_rating: bool
has_review_section: bool
has_product_gallery: bool
# listing
has_repeated_cards: bool
repeated_card_count: int
has_filter_panel: bool
has_sort_control: bool
has_pagination: bool
# editorial
has_author: bool
has_published_date: bool
has_modified_date: bool
has_article_body: bool
has_tags: bool
# community
has_question: bool
has_answer: bool
has_comments: bool
has_votes: bool
has_thread_structure: bool
has_faq_structure: bool
# knowledge/docs
has_code_blocks: bool
has_toc: bool
has_api_endpoint: bool
has_parameter_table: bool
has_version_info: bool
# corporate/legal
has_contact_info: bool
has_address: bool
has_policy_terms: bool
has_privacy_terms: bool
has_career_terms: bool
# transaction/protected
has_login_form: bool
has_password_field: bool
has_payment_fields: bool
has_captcha: bool
has_access_denied: bool
# graph
internal_link_count: int
external_link_count: int
product_link_count: int
category_link_count: int
profile_link_count: int
article_link_count: int
# text/layout
dominant_language: str | None
keyword_hits: dict[str, int]
```
---
## 6. Scoring 방식
단일 if-else가 아니라 evidence scoring으로 분류한다.
예:
```python
scores = {
"ProductDetailPage": 0.0,
"CategoryListingPage": 0.0,
"ArticlePage": 0.0,
"QAPage": 0.0,
"LoginPage": 0.0,
}
```
---
### 6.1 ProductDetailPage scoring 예시
```txt
schema.org Product +0.40
schema.org Offer +0.15
price detected +0.15
cart button +0.20
variant selector +0.15
SKU +0.10
product image gallery +0.10
review section +0.05
URL product hint +0.05
```
---
### 6.2 CategoryListingPage scoring 예시
```txt
repeated product cards +0.35
filter panel +0.20
sort control +0.15
pagination +0.10
many product links +0.20
breadcrumb category +0.10
URL category/list hint +0.05
```
---
### 6.3 ArticlePage scoring 예시
```txt
schema.org Article +0.35
schema.org NewsArticle +0.35
author +0.15
published date +0.15
article body +0.20
headline +0.10
tags +0.05
URL blog/news/article hint +0.05
```
---
### 6.4 QAPage scoring 예시
```txt
schema.org QAPage +0.35
question block +0.20
answer block +0.20
accepted answer +0.15
votes +0.10
comments +0.05
URL question/qna hint +0.05
```
---
### 6.5 LoginPage scoring 예시
```txt
password input +0.40
login keyword +0.20
email/user id input +0.15
submit button +0.10
signup/reset password links +0.10
```
---
## 7. Classification Result 모델
분류 결과는 단일 문자열이 아니라 아래 구조로 반환한다.
```python
@dataclass
class EvidenceItem:
key: str
value: str | int | float | bool | None
weight: float
source: str
message: str
@dataclass
class PageClassificationResult:
url: str
primary_page_type: str
secondary_page_types: list[str]
domain: str
archetype: str
main_entity_type: str | None
action_intents: list[str]
graph_roles: list[str]
confidence: float
alternatives: list[tuple[str, float]]
evidence: list[EvidenceItem]
should_analyze: bool
analyze_strategy: str
llm_policy: str
is_protected: bool
is_noise: bool
```
예:
```json
{
"primary_page_type": "CategoryListingPage",
"secondary_page_types": ["FilteredResultsPage"],
"domain": "Commerce",
"archetype": "Listing",
"main_entity_type": "Product",
"action_intents": ["Filter", "Compare", "Navigate"],
"graph_roles": ["CollectionHub", "RelationHub"],
"confidence": 0.88,
"should_analyze": true,
"analyze_strategy": "AnalyzeRelationsOnly",
"llm_policy": "RuleOnly",
"is_protected": false,
"is_noise": false
}
```
---
## 8. Analyze Strategy 정책
기존 `should_analyze_page(page_type, analyze_page_types)`는 유지하되, 내부를 확장한다.
기존 방식:
```txt
ProductPage -> analyze
CategoryPage -> skip
SearchPage -> skip
BoardPage -> skip
```
개선 방식:
```txt
ProductDetailPage -> AnalyzeFull
ArticlePage -> AnalyzeFull
BrandStoryPage -> AnalyzeFull
CategoryListingPage -> AnalyzeRelationsOnly
SearchResultsPage -> AnalyzeDiscoveryOnly
ForumThreadPage -> AnalyzeFull
ForumBoardPage -> AnalyzeRelationsOnly
FAQPage -> AnalyzeFull
QAPage -> AnalyzeFull
TermsPage -> AnalyzeDocumentOnly
PrivacyPolicyPage -> AnalyzeDocumentOnly
SitemapPage -> AnalyzeDiscoveryOnly
RobotsTxtPage -> AnalyzeMetadataOnly
LoginPage -> SkipProtected
CheckoutPage -> SkipProtected
PaymentPage -> SkipProtected
ErrorPage -> SkipNoise
NotFoundPage -> SkipNoise
```
---
## 9. LLM Policy 정책
LLM은 모든 페이지에 쓰지 않는다.
```txt
LLMFull
- ArticlePage
- ProductDetailPage
- BrandStoryPage
- ResearchPaperPage
- LegalArticlePage
LLMLight
- FAQPage
- QAPage
- DocumentationPage
- TutorialPage
LLMForAmbiguityOnly
- CategoryListingPage
- SearchResultsPage
- ForumBoardPage
- ArchivePage
- TagPage
RuleOnly
- SitemapPage
- RSSFeedPage
- RobotsTxtPage
- LoginPage
- CheckoutPage
- PaymentPage
Skip
- ErrorPage
- NotFoundPage
- AccessDeniedPage
- CaptchaPage
```
---
## 10. UnknownPage 처리
UnknownPage는 버리지 않는다.
분류 confidence가 낮은 경우:
```txt
primary_page_type = "UnknownPage"
graph_roles = ["UnknownPattern"]
analyze_strategy = "AnalyzeMetadataOnly"
llm_policy = "LLMForAmbiguityOnly" 또는 "NoLLM"
```
저장해야 할 정보:
```txt
url
title
text sample
html fingerprint
dom fingerprint
schema types
link pattern
button labels
forms
top keywords
embedding
classification alternatives
```
향후 clustering 대상:
```txt
UnknownPatternCluster
```
새 page type 후보 생성 예:
```txt
Unknown cluster 12
-> repeated comparison tables
-> product attributes
-> no cart button
-> many product links
=> ProductComparisonPage 후보
```
---
## 11. 구현 단계
## Phase 1. Taxonomy와 Result 모델 추가
### 작업
1. `page_type_taxonomy.py` 추가
2. PageDomain enum 추가
3. PageArchetype enum 추가
4. PageType enum 또는 문자열 상수 추가
5. EntityType enum 추가
6. ActionIntent enum 추가
7. GraphRole enum 추가
8. AnalyzeStrategy enum 추가
9. LLMPolicy enum 추가
10. `PageClassificationResult`, `EvidenceItem` dataclass 추가
### 완료 기준
- 기존 `ProductPage`, `CategoryPage` 등과 호환되어야 한다.
- 기존 코드에서 string page_type만 기대하는 부분은 깨지지 않도록 compatibility helper를 제공한다.
예:
```python
def get_legacy_page_type(result: PageClassificationResult) -> str:
return result.primary_page_type
```
---
## Phase 2. Signal Extractor 추가
### 작업
1. `page_signals.py` 추가
2. `page_signal_extractor.py` 추가
3. HTML에서 JSON-LD 추출
4. OpenGraph 추출
5. Twitter Card 추출
6. meta 태그 추출
7. button text 추출
8. form/input 추출
9. link pattern 추출
10. 반복 카드 후보 탐지
11. breadcrumb 후보 탐지
12. price/currency 후보 탐지
13. article author/date 후보 탐지
14. login/password/payment/captcha 후보 탐지
### 완료 기준
- 입력 HTML이 일부 깨져도 예외 없이 동작해야 한다.
- BeautifulSoup 또는 현재 프로젝트에서 사용하는 parser에 맞춰 구현한다.
- 없는 값은 빈 리스트/빈 dict/False로 처리한다.
---
## Phase 3. Scoring 기반 Page Type 분류
### 작업
1. `page_type_scorer.py` 추가
2. 주요 page type별 scoring function 작성
3. score normalize
4. top score와 alternatives 산출
5. confidence 계산
6. evidence 기록
최소 구현 대상:
```txt
ProductDetailPage
CategoryListingPage
SearchResultsPage
ArticlePage
BlogPostPage
QAPage
FAQPage
ForumBoardPage
ForumThreadPage
BrandStoryPage
AboutPage
ContactPage
DocumentationPage
APIReferencePage
JobPostingPage
PricingPage
LoginPage
CheckoutPage
TermsPage
PrivacyPolicyPage
SitemapPage
RSSFeedPage
ErrorPage
UnknownPage
```
### 완료 기준
- 단일 if-else return 금지
- 반드시 evidence list를 남긴다
- confidence가 낮으면 UnknownPage로 보낼 수 있어야 한다
---
## Phase 4. Analyze Strategy / LLM Policy 분리
### 작업
1. `page_analysis_policy.py` 추가
2. PageType -> AnalyzeStrategy mapping 작성
3. PageType -> LLMPolicy mapping 작성
4. 기존 `should_analyze_page()`를 compatibility 형태로 유지
5. 신규 함수 추가
예:
```python
def decide_analyze_strategy(result: PageClassificationResult) -> AnalyzeStrategy:
...
def decide_llm_policy(result: PageClassificationResult) -> LLMPolicy:
...
def should_analyze_page(result_or_page_type, analyze_page_types=None) -> bool:
...
```
### 완료 기준
- 기존 호출부가 바로 깨지지 않아야 한다.
- CategoryListingPage는 기본 skip이 아니라 `AnalyzeRelationsOnly`가 되어야 한다.
- SearchResultsPage는 `AnalyzeDiscoveryOnly`가 되어야 한다.
- Login/Checkout/Payment는 `SkipProtected`가 되어야 한다.
---
## Phase 5. 기존 Extractor 연결 수정
### 작업
기존 흐름:
```txt
classify_page(...)
-> page_type
-> should_analyze_page(page_type, analyze_page_types)
-> Extractor 실행
-> HybridExtractor 내부에서 LLM skip 판단
```
개선 흐름:
```txt
classify_page(...)
-> PageClassificationResult
-> decide_analyze_strategy(result)
-> decide_llm_policy(result)
-> strategy에 따라 Extractor 또는 relation/link extractor 실행
```
### 전략별 처리
```txt
AnalyzeFull
-> 기존 Extractor + HybridExtractor + LLM policy 적용
AnalyzeRelationsOnly
-> 링크, breadcrumb, category, repeated card 중심 추출
-> full LLM 금지
AnalyzeDiscoveryOnly
-> crawl candidate, result link, pagination만 추출
-> content claim 추출 금지
AnalyzeMetadataOnly
-> title, metadata, schema, canonical, link relation만 저장
AnalyzeDocumentOnly
-> 법률/정책/문서 구조 추출
-> 필요 시 LLM 사용
SkipProtected
-> 개인정보/계정/결제 페이지 분석 금지
SkipNoise
-> 저장 최소화 또는 제외
```
---
## Phase 6. Unknown Pattern 저장 기반 추가
### 작업
1. UnknownPage 또는 confidence 낮은 페이지를 별도 저장
2. DOM fingerprint 생성
3. text fingerprint 생성
4. link pattern summary 생성
5. 향후 clustering을 위한 embedding hook 추가
초기에는 clustering까지 구현하지 않아도 된다.
다만 데이터 구조는 남겨야 한다.
---
## Phase 7. 테스트 추가
### 단위 테스트 대상
```txt
ProductDetailPage
CategoryListingPage
SearchResultsPage
ArticlePage
QAPage
FAQPage
ForumThreadPage
DocumentationPage
JobPostingPage
LoginPage
CheckoutPage
TermsPage
SitemapPage
UnknownPage
```
### 테스트 샘플
각 page type에 대해 최소 HTML fixture를 만든다.
예:
```txt
tests/fixtures/pages/product_detail.html
tests/fixtures/pages/category_listing.html
tests/fixtures/pages/article.html
tests/fixtures/pages/qapage.html
tests/fixtures/pages/login.html
```
### 테스트 기준
```txt
primary_page_type이 기대값과 일치
confidence가 최소 기준 이상
evidence가 비어 있지 않음
analyze_strategy가 기대값과 일치
llm_policy가 기대값과 일치
protected page가 분석되지 않음
UnknownPage가 예외 없이 처리됨
```
---
## 12. 분류 예시
### 12.1 ProductDetailPage
입력 신호:
```txt
JSON-LD @type Product
price
add to cart
variant selector
product images
reviews
```
결과:
```json
{
"primary_page_type": "ProductDetailPage",
"domain": "Commerce",
"archetype": "Detail",
"main_entity_type": "Product",
"action_intents": ["Buy", "Review"],
"graph_roles": ["EntityAnchor"],
"confidence": 0.93,
"analyze_strategy": "AnalyzeFull",
"llm_policy": "LLMLight"
}
```
---
### 12.2 CategoryListingPage
입력 신호:
```txt
repeated product cards
filter panel
sort control
pagination
many product links
breadcrumb
```
결과:
```json
{
"primary_page_type": "CategoryListingPage",
"domain": "Commerce",
"archetype": "Listing",
"main_entity_type": "Product",
"action_intents": ["Filter", "Navigate", "Compare"],
"graph_roles": ["CollectionHub", "RelationHub"],
"confidence": 0.88,
"analyze_strategy": "AnalyzeRelationsOnly",
"llm_policy": "RuleOnly"
}
```
---
### 12.3 ArticlePage
입력 신호:
```txt
schema.org Article
headline
author
published date
article body
tags
```
결과:
```json
{
"primary_page_type": "ArticlePage",
"domain": "Editorial",
"archetype": "Article",
"main_entity_type": "Article",
"action_intents": ["Read"],
"graph_roles": ["ClaimSource", "EntityAnchor"],
"confidence": 0.91,
"analyze_strategy": "AnalyzeFull",
"llm_policy": "LLMFull"
}
```
---
### 12.4 LoginPage
입력 신호:
```txt
password input
email input
login button
reset password link
```
결과:
```json
{
"primary_page_type": "LoginPage",
"domain": "Transaction",
"archetype": "Form",
"main_entity_type": null,
"action_intents": ["Login"],
"graph_roles": ["TransactionOnly"],
"confidence": 0.96,
"analyze_strategy": "SkipProtected",
"llm_policy": "Skip"
}
```
---
## 13. 하위 호환성 요구
기존 코드가 아래처럼 page_type 문자열을 기대할 수 있다.
```python
page_type = classify_page(...)
should_analyze_page(page_type, analyze_page_types)
```
따라서 처음부터 모든 호출부를 바꾸지 말고, compatibility layer를 둔다.
권장:
```python
def classify_page_legacy(*args, **kwargs) -> str:
result = classify_page(*args, **kwargs)
return result.primary_page_type
```
또는:
```python
def normalize_page_type(page_type_or_result) -> str:
if isinstance(page_type_or_result, PageClassificationResult):
return page_type_or_result.primary_page_type
return str(page_type_or_result)
```
---
## 14. 기존 page type과 신규 page type 매핑
```txt
ProductPage -> ProductDetailPage
CategoryPage -> CategoryListingPage 또는 CategoryPage
SearchPage -> SearchResultsPage
BoardPage -> ForumBoardPage
NoticePage -> PublicNoticePage 또는 NoticePage
BrandStoryPage -> BrandStoryPage
PromotionPage -> PromotionPage 또는 CampaignLandingPage
UnknownPage -> UnknownPage
ReviewPage -> ReviewPage 또는 ProductReviewPage
```
기존 이름은 당분간 alias로 유지한다.
---
## 15. 코딩 원칙
1. 단일 if-return 방식으로 확장하지 말 것
2. evidence를 반드시 남길 것
3. confidence를 반드시 계산할 것
4. PageType 하나만 반환하지 말 것
5. Category/Search/Board 계열을 무조건 skip하지 말 것
6. Login/Checkout/Payment는 보호 페이지로 처리할 것
7. UnknownPage는 폐기하지 말고 저장 가능한 구조로 만들 것
8. JSON-LD, OpenGraph, meta, DOM, link graph를 모두 signal로 사용할 것
9. 다국어 키워드 확장을 고려할 것
10. 기존 호출부가 깨지지 않도록 compatibility helper를 제공할 것
---
## 16. 최소 완료 기준
이번 작업의 최소 완료 기준은 다음과 같다.
```txt
1. PageClassificationResult dataclass 추가
2. PageDomain / PageArchetype / PageType / EntityType / ActionIntent / GraphRole / AnalyzeStrategy / LLMPolicy 정의
3. JSON-LD / OpenGraph / URL / DOM / text 기반 signal 추출
4. 최소 20개 page type scoring 구현
5. CategoryListingPage가 AnalyzeRelationsOnly로 처리됨
6. SearchResultsPage가 AnalyzeDiscoveryOnly로 처리됨
7. LoginPage / CheckoutPage / PaymentPage가 SkipProtected로 처리됨
8. ProductDetailPage / ArticlePage / FAQPage / QAPage는 AnalyzeFull 가능
9. UnknownPage가 evidence와 함께 반환됨
10. 기존 should_analyze_page 호환성 유지
11. 테스트 fixture 10개 이상 추가
12. pytest 통과
```
---
## 17. 최종 목표
이 작업의 최종 목표는 `page_classifier.py`를 다음 수준으로 확장하는 것이다.
기존:
```txt
URL/text 기반 page type 분류기
```
개선:
```txt
Semantic Page Understanding Layer
```
최종 파이프라인:
```txt
Raw Page
-> Signal Extraction
-> Evidence Scoring
-> Domain Classification
-> Archetype Classification
-> Main Entity Classification
-> Action Intent Classification
-> Graph Role Assignment
-> Analyze Strategy Decision
-> LLM Policy Decision
-> Unknown Pattern Storage
```
이렇게 해야 온톨로지 플랫폼이 특정 쇼핑몰이나 특정 사이트에 종속되지 않고, 인터넷 전체의 다양한 페이지를 의미 단위로 해석할 수 있다.
---
## 18. Codex 작업 요청 요약
Codex는 이 문서를 기준으로 다음 작업을 수행한다.
```txt
1. 현재 page_classifier.py를 확인한다.
2. 기존 호출부와 테스트를 확인한다.
3. taxonomy / signal / scoring / policy 레이어를 분리한다.
4. 기존 단순 page_type 문자열 반환 구조를 PageClassificationResult 중심으로 확장한다.
5. 기존 코드가 깨지지 않도록 legacy compatibility를 유지한다.
6. Category/Search/Board를 단순 skip하지 않고 strategy 기반으로 처리한다.
7. protected/transaction page는 안전하게 skip한다.
8. UnknownPage는 evidence와 함께 저장 가능하게 만든다.
9. 최소 fixture 테스트를 추가한다.
10. pytest로 회귀 테스트를 확인한다.
```
---
## 19. 주의 사항
- 기존 엔진을 대규모로 폐기하지 말 것.
- 현재 parser, crawler, extractor 흐름을 먼저 파악한 뒤 최소 침습 방식으로 확장할 것.
- page type enum 확장은 허용하되, extractor 전체를 한 번에 갈아엎지 말 것.
- LLM 사용량이 늘어나지 않도록 `LLMPolicy`를 반드시 적용할 것.
- protected page에서 개인정보나 계정 정보를 추출하지 말 것.
- 분류가 애매할 경우 억지로 하나의 타입에 넣지 말고 alternatives와 confidence를 남길 것.
- Evidence 기반 디버깅이 가능해야 한다.
---
## 20. 향후 확장 방향
이번 작업 이후 다음 단계로 확장할 수 있다.
```txt
1. UnknownPage clustering
2. DOM fingerprint 기반 template detection
3. site-specific learned page archetype
4. screenshot 기반 visual block classification
5. multilingual keyword dictionary
6. schema.org type mapping 강화
7. page type별 extraction schema 자동 선택
8. crawl priority와 page type 연동
9. graph relation confidence와 page evidence 연동
10. admin UI에서 page classification 결과 검토
```