Generative AI and automation
An AI web agent learns to build scrapers
AutoScraper uses a language model to generate instructions for collecting webpage data. Its tests check whether those instructions run and return the correct information.
A web scraper can save repeated manual copying, but a script built for one page may fail when a site changes. AutoScraper uses a language model to generate scraper actions and tests whether those actions can actually run.

From page structure to an executable scraper
The paper introduces a task in which a language model produces action sequences to extract information from vertical web pages. AutoScraper has two stages. Progressive generation uses the hierarchy of a page's HTML, the code structure that describes webpage elements, to help the model understand long documents in smaller steps. Synthesis then combines information from multiple related pages into a more complete action sequence. The method is designed to use similarities across pages while adapting to different layouts.
The authors also propose an executability metric. Conventional extraction scores such as precision and recall measure the correctness of extracted items. They can overlook cases where the generated action sequence cannot run at all. AutoScraper therefore assesses whether a scraper executes and whether it returns correct information, alongside information-extraction measures.
Experiments compare the method with chain-of-thought and Reflexion baselines across three datasets: SWDE, EXTEND SWDE and DS1. Table 2 (PDF p. 6) reports execution and extraction measures. The authors find that AutoScraper improves the proportion of correct, executable action sequences and lowers the proportion that cannot execute in their tested settings. They report that larger language models are more stable at understanding page structure and reflecting on execution results. Table 3 (PDF p. 7) studies the method's components.
The paper focuses on information extraction from vertical webpages. That means pages organized around a category, such as products or properties, rather than unrestricted web navigation. The framework depends on the underlying model's ability to understand HTML.
A metric that notices failures before extraction
The distinction between an executable workflow and a correct output is important for automation. A scraper that fails to run produces no useful data, even if its intended extraction logic appears sound. Measuring execution can therefore reveal a failure mode that item-level accuracy alone misses. This is a methodological contribution as well as a system design.
The authors' limitations discussion (PDF p. 9) notes that AutoScraper is restricted to vertical web information-extraction tasks and may not transfer to broader interactive benchmarks. Results depend on the language model and the tested webpages. The study does not measure maintenance cost, website-policy compliance or the effects of live website changes over time.
Possible implications for practice
The results suggest that the reliability of data automation may be better understood when teams track whether a generated procedure runs as intended, in addition to whether its output looks correct. This framing could inform evaluations of AI-assisted data collection, while leaving governance, permissions and ongoing maintenance to separate review.
Bibliography & sources
- Wenhao Huang; Zhouhong Gu; Chenghao Peng; Zhixu Li; Jiaqing Liang; Yanghua Xiao; Liqian Wen; Zulong Chen (2024). “AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation.” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 2371–2389. DOI: 10.18653/v1/2024.emnlp-main.141. Source paper
