Description review
Big Data Developer
PradeepIT Consulting Services Pvt Ltd · India · back to the listing
HR standards
44/100
poor
Title ↔ description
83/100
solid
Reads as
Data Engineer
100% confident
How others title the same work
Large employers
- Data Engineer, Product Analytics Meta
- Senior Staff Data Engineer Mozilla
- Staff, Analytics Engineer, GTM Data Science & Analytics Twilio
- Senior Software Engineer - Data Platform Coinbase
- Senior Security Data Engineer Doordash
Startups
- Business Intelligence Engineer GoSats
- IPinfo.io | Data Engineer | REMOTE (Anywhere) | Full-time IPinfo.io
- Senior Data Engineer Camber
- Senior Data Engineer Instrumentl
- Cardog | Toronto, Canada or REMOTE | Full-time Cardog
What the listing never says
- No section describes what the person would actually do. Scope clarity
- No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency
- No location or timezone policy stated, so a candidate cannot tell where they may work from. Scope clarity
The listing, marked up
Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.
L3 Big datadeveloper ( 7-10 Years)
1. Design, develop, and implement highly scalable and distributed big datasolutions using Hadoop ecosystem technologies such as HBase, Hive, Kudu, andSpark.
2. Architect HBase schemas and data models to accommodate evolving businessrequirements and ensure optimal performance for data storage and retrievaloperations.
3. Develop complex Hive queries and data processing pipelines to transform rawdata into structured formats suitable for analysis and reporting.
4. Implement data ingestion pipelines using Spark Streaming and Spark SQL forreal-time processing of streaming data sources, ensuring high throughput andlow latency.
5. Optimize Spark applications for performance and resource utilization,including tuning RDD transformations, optimizing data partitioning strategies,and leveraging in-memory caching.
6. Utilize advanced features of Spark MLlib for machine learning tasks such asclassification, regression, clustering, and collaborative filtering.
7. Design and deploy Kudu tables for fast analytical queries and real-timeanalytics, leveraging Kudus unique combination of fast analytics and fast dataingestion.
8. Collaborate with data scientists to integrate machine learning models intoSpark workflows and productionize them for real-time predictions and analytics.
9. Troubleshoot performance bottlenecks, data quality issues, and systemfailures in big data applications and infrastructure, and implement solutionsto address them.
10. Stay abreast of emerging technologies and best practices in big data processingand analytics, and evaluate their potential impact on our architecture andsolutions.
Originally posted on Himalayas
1. Design, develop, and implement highly scalable and distributed big datasolutions using Hadoop ecosystem technologies such as HBase, Hive, Kudu, andSpark.
2. Architect HBase schemas and data models to accommodate evolving businessrequirements and ensure optimal performance for data storage and retrievaloperations.
3. Develop complex Hive queries and data processing pipelines to transform rawdata into structured formats suitable for analysis and reporting.
4. Implement data ingestion pipelines using Spark Streaming and Spark SQL forreal-time processing of streaming data sources, ensuring high throughput andlow latency.
5. Optimize Spark applications for performance and resource utilization,including tuning RDD transformations, optimizing data partitioning strategies,and leveraging in-memory caching.
6. Utilize advanced features of Spark MLlib for machine learning tasks such asclassification, regression, clustering, and collaborative filtering.
7. Design and deploy Kudu tables for fast analytical queries and real-timeanalytics, leveraging Kudus unique combination of fast analytics and fast dataingestion.
8. Collaborate with data scientists to integrate machine learning models intoSpark workflows and productionize them for real-time predictions and analytics.
9. Troubleshoot performance bottlenecks, data quality issues, and systemfailures in big data applications and infrastructure, and implement solutionsto address them.
10. Stay abreast of emerging technologies and best practices in big data processingand analytics, and evaluate their potential impact on our architecture andsolutions.
Originally posted on Himalayas