Government websites as data: a methodological pipeline with application to the websites of municipalities in the United States

The content of a government’s website is an important source of information about policy priorities, procedures, and services. Existing research on government websites has relied on manual methods of website content collection and processing, which imposes cost limitations on the scale of website data collection. In this research note, we propose that the automated collection of website content from large samples of government websites can offer relief from the costs of manual collection, and enable contributions through large-scale comparative analyses. We also provide software to ease the use of this data collection method. In an illustrative application, we collect textual content from the websites of over two hundred municipal governments in the United States, and study how website content is associated with mayoral partisanship. Using statistical topic modeling, we find that the partisanship of the mayor predicts differences in the contents of city websites that align with differences in the platforms of Democrats and Republicans. The application illustrates the utility of website content data extracted via our methodological pipeline.

This is an Accepted Manuscript of an article published by Taylor & Francis in Journal of Information Technology & Politics on 2022-10-02, available online: https://www.tandfonline.com/10.1080/19331681.2021.1999880.

Files

Metadata

Work Title Government websites as data: a methodological pipeline with application to the websites of municipalities in the United States
Access
Open Access
Creators
  1. Markus Neumann
  2. Fridolin Linder
  3. Bruce Desmarais
Keyword
  1. Webscraping
  2. Government websites
  3. Text analysis
  4. Cities
  5. Mayors
  6. Topic models
License CC BY-NC 4.0 (Attribution-NonCommercial)
Work Type Article
Publisher
  1. Journal of Information Technology and Politics
Publication Date November 24, 2021
Publisher Identifier (DOI)
  1. https://doi.org/10.1080/19331681.2021.1999880
Deposited January 23, 2023

Versions

Analytics

Collections

This resource is currently not in any collection.

Work History

Version 1
published

  • Created
  • Added manuscript_jitp_rr2.pdf
  • Added Creator Markus Neumann
  • Added Creator Fridolin Linder
  • Added Creator Bruce Desmarais
  • Published
  • Updated Keyword, Subtitle, Publication Date Show Changes
    Keyword
    • Webscraping, Government websites, Text analysis, Cities, Mayors, Topic models
    Subtitle
    • a methodological pipeline with application to the websites of municipalities in the United States
    Publication Date
    • 2021-01-01
    • 2021-11-24
  • Updated