Web Scraping Specialist
You will extract data from complex websites using Python or JavaScript and tools such as BeautifulSoup Scrapy or Selenium. You will build reliable scraping workflows handle pagination and dynamic AJAX content clean and format data manage NoSQL databases and monitor processes to resolve issues. You may also deploy scraping jobs using cloud services and apply machine learning for data cleaning categorization or predictive analysis.
Responsibilities
- Write test and refine data extraction code
- Retrieve data from online sources including paginated and AJAX-loaded content
- Clean and format extracted data
- Store and manage scraped data in databases
- Optimize database access speed and data integrity
- Monitor scraping processes and resolve issues
Requirements
- Demonstrated ability to extract data from complex websites with minimal supervision
- Portfolio or examples of past web scraping projects
- Proficiency in Python or JavaScript
- Strong skills with BeautifulSoup Scrapy or Selenium
- Knowledge of asynchronous programming multithreading and distributed scraping
- In-depth knowledge of HTML CSS JavaScript and the Document Object Model
- Experience with MongoDB or Cassandra
- Ability to design efficient storage solutions and manage data integrity
- Ability to apply machine learning algorithms for data cleaning categorization or predictive analysis
- Experience with AWS Google Cloud or Azure
- Active participation in relevant open-source projects
Benefits
- Benefits package
- Equity package