Key Responsibilities:
- Develop and maintain robust web scraping solutions to extract structured and unstructured data from US county property appraiser/assessor websites and public records and property-related data sources.
- Handle complex scraping challenges such as CAPTCHA, session handling, rate limiting, agile content (JavaScript-heavy sites).
- Design and build end-to-end data pipelines: Data extraction transformation mapping loading (ETL/ELT).
- Perform data transformation and attribute mapping to align with internal data models.
- Develop automated workflows/applications for periodic database updates, data refresh and validation, and ad-hoc data extraction requirements.
- Write and optimize SQL/PostgreSQL queries for data ingestion and data updates and maintenance.
- Ensure data quality, consistency, and integrity across datasets.
- Collaborate with analytics, GIS, and business teams to support data-driven use cases.
Requirements
- Experience: 2-3 years of hands-on experience in Python development.
- Strong expertise in web scraping using Python libraries,
such as Beautiful Soup, Scrapy, Selenium, Playwright.
- Experience handling anti-scraping mechanisms (CAPTCHA, headers, proxies, etc.).
- Solid experience with PostgreSQL / SQL: Writing complex queries, database updates, joins, indexing, performance tuning.
- Proven experience in building data pipelines (ETL/ELT).
- Strong understanding of data structures and large dataset handling.
- Experience in automation scripting and workflow development.
- Good understanding of data transformation and mapping techniques.
- Strong problem-solving ability and attention to detail.
Good to Have (Preferred Skills):
- Experience with geospatial data / GIS tools: GeoPandas, PostGIS, QGIS, Shapely.
- Exposure to US Property Casualty Insurance domain.
- Understanding of property-level datasets and attributes.
- Experience with data validation frameworks and logging/monitoring.
- Familiarity with data visualisation or an