Skip to content

Don't parse an HTTP error page's body as feed content - #597

Open
hikmetba-bit wants to merge 1 commit into
kurtmckee:mainfrom
hikmetba-bit:fix/460-dont-parse-http-error-body
Open

hikmetba-bit wants to merge 1 commit into
kurtmckee:mainfrom
hikmetba-bit:fix/460-dont-parse-http-error-body

Conversation

@hikmetba-bit

Copy link
Copy Markdown

Summary

Fixes #460.

>>> feedparser.parse("http://httpstat.us/500")
{'bozo': 1, 'entries': [], 'feed': {}, ...,
 'bozo_exception': SAXParseException('syntax error'), 'status': 500, ...}

http.get() returned the HTTP response body regardless of status code, so a server's 500 error page (HTML, plain text, whatever the error handler produces) got fed straight into the XML parser. That produces a bozo_exception that looks like a malformed feed, when the real problem — the HTTP failure — is already recorded cleanly in result["status"] a few lines above.

Fix: in feedparser/http.py, get() now returns b"" for any response with status_code >= 400, the same empty-body return it already uses when the request itself raises RequestException. status/headers/href are still set from the real response before that check, so callers can still see exactly what happened (status == 500), they just don't also get a misleading parse error bozo_exception.

Test plan

  • Added tests/test_http.py with two tests using the responses mock already used elsewhere in the suite (tests/helpers.py):
    • test_error_status_body_is_not_parsed: a mocked 500 response with a plain-text body → status == 500, bozo is falsy, no bozo_exception key, entries == [].
    • test_success_status_body_is_still_parsed: a mocked 200 RSS response, to confirm normal parsing is unaffected.
  • Full suite: pytest tests/ → 4299 passed, 8 skipped (was 4297 passed, 8 skipped on main before this change — the +2 are the new tests, zero regressions).
  • mypy feedparser/http.py → no issues. flake8 on both changed files → clean.

Co-Authored-By: Claude Sonnet 5 [email protected]
🤖 Generated with Claude Code

feedparser.parse('http://httpstat.us/500') fed the 500 error page's
body straight to the XML parser, producing a confusing
bozo_exception (e.g. SAXParseException('syntax error')) that looks
like a malformed feed. The HTTP failure is already recorded cleanly
in result['status'], set unconditionally a few lines above.

http.get() now returns an empty body for any response with
status_code >= 400, matching the existing empty-body return already
used for a RequestException. bozo/bozo_exception are then left
unset by the empty-content parse, same as parsing b"".

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feedparser should not attempt to parse HTTP error pages

1 participant