Senior Member, Data Engineer (Data Extraction)
D. E. Shaw
D. E. Shaw
We are looking for resourceful and exceptional candidates for the Data Engineer role within our product development teams based out of Hyderabad, Bengaluru and Gurugram. At DESIS, the Data Engineers develop Web Robots, or Web Spiders, that crawl through the web and retrieve data in the form of HTML, plain text, PDFs, Excel, and any other format that is either structured or unstructured. The job functions of the engineer also include scraping the website data into a structured format and building automated and custom reports on the downloaded data that are used as knowledge for business purposes. The team also works on automating end-to-end data pipelines.
Responsibilities:
As a member of the Data Engineering team, you will be responsible for various aspects of data extraction, such as understanding the data requirements of the business group, reverse-engineering the website, its technology, and the data retrieval process, re-engineering by developing web robots to automate the extraction of the data, and building monitoring systems to ensure the integrity and quality of the extracted data. You will also be responsible for managing the changes to the website's dynamics and layout to ensure clean downloads, building scraping and parsing systems to transform raw data into a structured form, and offering operations support to ensure high availability and zero data losses. Additionally, you will be involved in other tasks such as storing the extracted data in the recommended databases, building high-performing, scalable data extraction systems, and automating data pipelines.
Who we're looking for:
The ideal candidate should hold-
Basic qualifications:
2 to 4 years of experience in website data extraction and scraping
Good knowledge of relational databases, writing complex queries in SQL, and dealing with ETL operations on databases
Proficiency in Python for performing operations on data
Expertise in Python frameworks like Requests, UrlLib2, Selenium, Beautiful Soup, and Scrapy
A good understanding of HTTP requests and responses, HTML, CSS, XML, JSON, and JavaScript
Expertise with debugging tools in Chrome to reverse engineer website dynamics
A good academic background and accomplishments
A BCA/MCA/BS/MS degree with a good foundation and practical application of knowledge in data structures and algorithms
Problem-solving and analytical skills
Good debugging skills
Ready to apply?
Sign up first — takes a minute — and you get Hiro’s take on this role, a resume tailored to it, and (if available) a referral from a real employee at D. E. Shaw. All free with your Pro gift.
Sign up to apply