Repository navigation
⚡️ Speed up method ElementHtml._get_children_html by 234% - #4087
Merged
mpolomdeepsense merged 10 commits intoSep 10, 2025
Conversation
Here is a **faster rewrite** of your program, based on your line profiling results, the imported code constraints, and the code logic. ### Key optimizations. - **Avoid repeated parsing:** The hotspot is in recursive calls to `child.get_html_element(**kwargs)`, each of which is re-creating a new `BeautifulSoup` object in every call. Solution: **Pass down and reuse a single `BeautifulSoup` instance** when building child HTML elements. - **Minimize object creation:** Create `soup` once at the *topmost* call and reuse for all children and subchildren. - **Reduce .get_text_as_html use:** Optimize to only use the soup instance when really necessary and avoid repeated blank parses. - **Avoid double wrapping:** Only allocate wrappers and new tags if absolutely required. - **General micro-optimizations:** Use `None` instead of `or []`, fast-path checks on empty children, etc. - **Preserve all comments and signatures as specified.** Below is the optimized version. ### Explanation of improvements - **Soup passing**: The `get_html_element` method now optionally receives a `_soup` kwarg. At the top of the tree, it is `None`, so a new one is created. Then, for all descendants, the same `soup` instance is passed via `_soup`, avoiding repeated parsing and allocation. - **Children check**: `self.children` is checked once, and the attribute itself is kept as a list (not or-ed with empty list at every call). - **No unnecessary soup parsing**: `get_text_as_html()` doesn't need a soup argument, since it only returns a Tag (from the parent module). - **No changes to existing comments, new comments added only where logic was changed.** - **Behavior (output and signature) preserved.** This **avoids creating thousands of BeautifulSoup objects recursively**, which was the primary bottleneck found in the profiler. The result is vastly improved performance, especially for large/complex trees.
mpolomdeepsense
requested changes
Sep 3, 2025
mpolomdeepsense
left a comment
Contributor
There was a problem hiding this comment.
Looks good but CI is throwing error because there was no update to changelog. Pls fix.
Contributor
Author
|
@mpolomdeepsense I have modified the changelog, it should be ready to merge. I noticed there seems to be an issue in |
Contributor
Author
|
@mpolomdeepsense ready to merge |
mpolomdeepsense
approved these changes
Sep 10, 2025
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
📄 234% (2.34x) speedup for
ElementHtml._get_children_htmlinunstructured/partition/html/convert.py⏱️ Runtime :
12.3 milliseconds→3.69 milliseconds(best of101runs)📝 Explanation and details
Here is a faster rewrite of your program, based on your line profiling results, the imported code constraints, and the code logic.
Key optimizations.
child.get_html_element(**kwargs), each of which is re-creating a newBeautifulSoupobject in every call.Solution: Pass down and reuse a single
BeautifulSoupinstance when building child HTML elements.souponce at the topmost call and reuse for all children and subchildren.Noneinstead ofor [], fast-path checks on empty children, etc.Below is the optimized version.
Explanation of improvements
get_html_elementmethod now optionally receives a_soupkwarg. At the top of the tree, it isNone, so a new one is created. Then, for all descendants, the samesoupinstance is passed via_soup, avoiding repeated parsing and allocation.self.childrenis checked once, and the attribute itself is kept as a list (not or-ed with empty list at every call).get_text_as_html()doesn't need a soup argument, since it only returns a Tag (from the parent module).This avoids creating thousands of BeautifulSoup objects recursively, which was the primary bottleneck found in the profiler. The result is vastly improved performance, especially for large/complex trees.
✅ Correctness verification report:
🌀 Generated Regression Tests and Runtime
To edit these changes
git checkout codeflash/optimize-ElementHtml._get_children_html-mcsd67coand push.