The situation
Crawl4AI is an open-source tool for crawling websites and extracting data, widely used to feed web content into AI systems. Many developers want to run it as a server, inside a Docker container, and call it through an API rather than as a Python library. That server side needed more capabilities, and some bugs were affecting how crawls behaved.
What I contributed
- A streaming endpoint, so results arrive as each page is crawled instead of after the whole job finishes.
- A utilization endpoint, so you can see how busy the server is.
- Bug fixes in the Docker and server code affecting crawl behaviour.
- SDK/API parity work, so the API offers the same options as the Python library.
- Tests covering the new behaviour.
Why it matters
Most of my client work is private, so I can’t show the code. This is different: it’s public, reviewed by the project’s maintainers, and merged into a tool other developers rely on.
You can check it yourself: my pull requests to Crawl4AI and the Crawl4AI repository.
What this means for you
If you need data from websites — prices, listings, documents, directories — I can build a crawler that runs on its own and exposes the results through your own API.