Repository navigation
feat: 실거래가 원본 데이터를 S3에 적재하는 수집 파이프라인 추가 - #4
Conversation
기존에는 로컬 scripts/results 디렉터리의 JSON 파일을 읽어 DB에 적재하는 스텝만 존재했다. 공공데이터포털 API에서 직접 실거래가를 수집하고 법정동 코드를 조회하는 단계를 추가해 법정동 코드 갱신 -> 수집(API -> S3) -> 적재(S3 -> DB) 전체 파이프라인으로 확장하고, main.py에 흩어져 있던 로직을 app/scheduler 패키지로 분리했다. 설정값은 app/config.py의 pydantic-settings로 통일해 관리한다. #2 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017vjm8qqWBxUGvzb5kkGQfu
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
- get_object/put_object/list_objects_v2 호출에 ExpectedBucketOwner를 추가해 의도치 않은 버킷에 접근하는 것을 방지 - collect_sigungu_codes의 행 필터링/누적 로직을 별도 함수로 분리해 인지 복잡도를 허용 범위 이내로 낮춤 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017vjm8qqWBxUGvzb5kkGQfu
|



Summary
법정동 코드 갱신 -> 수집(API -> S3) -> 적재(S3 -> DB)3단계로 확장 (run_scheduled_pipeline)main.py에 있던 수집/변환/적재 로직을app/scheduler패키지(collect.py,legal_dong.py,transform.py,ingest.py,jobs.py)로 분리app/config.py의 pydantic-settings(Settings)로 통일 (DATABASE_URL,S3_BUCKET,S3_PREFIX,DATA_GO_KR_SERVICE_KEY),.env.example추가boto3,defusedxml,requests,pydantic-settings의존성 추가Test plan
ruff format .,ruff check .통과run_scheduled_pipeline수동 실행으로 S3 업로드/DB 적재 확인Closes #2
🤖 Generated with Claude Code
https://claude.ai/code/session_017vjm8qqWBxUGvzb5kkGQfu