Google Blogger Backup & Data Extraction Project Summary
클라우드 개발 환경(GitHub Codespaces) 내에서 구글 블로거(Blogger/Blogspot)에 축적된 115개의 포스트 데이터를 로컬 디렉터리로 이식하기 위해 총 3가지의 시나리오를 설계하고 검증을 진행했습니다. 구글의 철저한 보안 메커니즘และ 호스트 차단 정책 속에서 각 방식이 직면한 기술적 한계와 최종 성공 경로에 대한 정밀 분석 보고서입니다.
-
1. 비인증 공개 RSS/Atom 피드 크롤러 방식 실패 (Blocked)
깃허브 오픈소스 리포지토리의 기본 크롤링 메커니즘을 이식하여 세팅했습니다. 인증 없이 블로그 식별자 주소 뒤에 아톰 피드 주소를 덧붙여 간편하게 XML 데이터를 요청하는 방식입니다. 로컬 PC 환경에서는 일시적으로 작동할 수 있으나, 가상 클라우드 인프라(Microsoft Azure/GitHub 데이터 센터) 환경에서 구글 서버로 접근을 찌르는 순간 구글 보안 엔진이 이를 악성 트래픽(DDoS 또는 크롤링 봇)으로 인지했습니다. 그 결과 데이터 반환 대신 일반 로그인 세션 주소인
://blogger.com로 강제 리다이렉션을 발생시켜too many redirects (MaxRetryError)장벽에 가로막혔습니다. -
2. 구글 Cloud Console 공식 API 연동 방식 (OAuth 2.0) 실패 (Scope Mismatch)
구글이 합법적으로 승인하는 규격에 맞추어 Cloud Console에서 프로젝트를 생성하고, 데스크톱 인증 열쇠 파일(
client_secret.json)을 발급받아 환경을 구축했습니다. 아이디와 비밀번호 유출 없이 비공개 드래프트 글까지 실시간 원격 동기화할 수 있는 가장 정석적인 루트입니다. 그러나 클라우드 에이전트 터미널 환경 특성상 기존 실습 로그에 남아있던 오염된 세션 환경 변수(BLOGGER_NAME등)와 내부 가상환경 캐시 파일(token.json)이 파이썬 실행 레이어에서 꼬이는 현상이 발생했습니다. 소스코드를 올바르게 수정하더라도 캐시된 메모리가 우선 호출되면서 구글 인증 서버에 불완전한 스코프 주소(https://googleapis.com)를 전송하게 되었고, 이로 인해 구글 보안 아키텍처가Error 400: invalid_scope오류를 뿜으며 진입을 철저히 차단했습니다. -
3. 구글 테이크아웃 피드 포팅 및 데이터 파싱 방식 성공 (Success)
네트워크 레이어의 차단 장벽을 완벽하게 무력화하기 위해 구글 계정 소유자가 직접 데이터를 내보내는
Google Takeout서비스를 채택했습니다. 다운로드한 백업 패키지 내부에서 블로그 레이아웃 테마(XML) 파일 대신, 실제 115개의 본문 원본 텍스트 데이터가 집약된feed.atom파일을 확보했습니다. 이를 코드스페이스 로컬 환경에 업로드한 후 BeautifulSoup 라이브러리의lxml파서 엔진을 활용해 파이썬 스크립트로 내부 XML 트리 구조를 정밀 분석했습니다. 설정 및 디자인 데이터와 실제 포스트 엔트리 데이터를 완벽하게 구별해 냄으로써, 단 1초 만에 323KB 용량의 온전한 텍스트 데이터베이스(backup.json)를 마이그레이션하는 데 완벽히 성공했습니다.
최종 수집된 텍스트 자산을 기반으로 파일 저장 규칙(YYYY-MM-DD-title.md) 및 Jekyll 스타일의 YAML 프론트매터(Front Matter) 규칙을 적용하는 추가 스크립트를 구동하여, 115개의 독립된 정적 블로그용 마크다운 파일로 분할 적재하는 영구 아카이빙 파이프라인까지 완벽하게 완수하였습니다.
Within a cloud development environment (GitHub Codespaces), we designed and validated three technical scenarios to migrate 115 accumulated post data from Google Blogger (Blogger/Blogspot) into a local repository. This is a comprehensive technical report analyzing the security barriers encountered and the ultimate successful path established through this exploration.
-
1. Unauthenticated Public RSS/Atom Feed Crawler Method Blocked
This scenario utilized an open-source parsing mechanism to query raw XML feed data by appending the atom feed path to the blog identifier. While functional on residential network connections, executing these scripts inside cloud infrastructure (Microsoft Azure / GitHub data centers) immediately triggered Google's automated threat detection. Recognizing the data center IP as an automated scraper bot, Google's server refused to serve the data payload and instead forced a loop of security redirects toward the standard login gate (
://blogger.com), generating a terminaltoo many redirects (MaxRetryError)failure. -
2. Google Cloud Console Official API Integration (OAuth 2.0) Scope Mismatch
To implement a secure, authorized synchronization pipeline, we instantiated a project on the Google Cloud Console and generated a desktop application client configuration file (
client_secret.json). This approach provides a secure handshake to fetch both public content and private drafts without exposing raw credentials. However, the headless cloud terminal environment suffered from session environment variable pollution (such asBLOGGER_NAME) and local cache synchronization conflicts (token.json). Even when the code was structurally accurate, the underlying Python authentication library prioritized cached configurations, transmitting an incomplete scope string (https://googleapis.com) to the authentication server, which ultimately raised an immutableError 400: invalid_scopebarrier. -
3. Google Takeout Feed Porting and Core Parsing Method Success
To fully bypass network-level friction and API bottlenecks, we pivoted to leveraging Google's official data extraction service, Google Takeout. Within the exported package, we bypassed the secondary layout schema (XML) files and successfully targeted the raw
feed.atomfile, which contained the entire historical text payload of the 115 posts. After porting this file into the Codespace storage layer, a custom Python script equipped with the high-performancelxmlparser engine was executed. The script successfully decoupled presentation parameters from structural content nodes, instantly extracting a clean 323KB JSON text database (backup.json) without any credential exposure or protocol blocks.
Utilizing the extracted core assets, an automated downstream formatting script was compiled to enforce structured file naming conventions (YYYY-MM-DD-title.md) and generate Jekyll-compliant YAML Front Matter blocks. This successfully transformed the centralized ledger into 115 standalone Markdown (.md) documents, establishing a permanent and autonomous static blog baseline.
댓글
댓글 쓰기