Description review
Data Engineer (Architect)
George Bernard Consulting · Sri Lanka · back to the listing
HR standards
35/100
poor
Title ↔ description
84/100
solid
Reads as
Data Engineer
100% confident
How others title the same work
Large employers
- Senior Software Engineer - Data Platform Coinbase
- Senior Security Data Engineer Doordash
- Data Engineer, Product Analytics Meta
- Senior Staff Data Engineer Mozilla
- Staff, Analytics Engineer, GTM Data Science & Analytics Twilio
Startups
- Senior Data Engineer Camber
- Cardog | Toronto, Canada or REMOTE | Full-time Cardog
- Connie Health | Senior Analytics Engineer | Boston, MA | HYBRID | $150k-$185k + equity Connie Health
- DoubleVerify (DV Scibids) | Sr. Data Engineer I | ONSITE/HYBRID - New York, NY (3x/week) | Full-time | $89K–$178K DoubleVerify (DV Scibids)
- Forecasting Research Institute (FRI) | Data Engineer | REMOTE | Full-time Forecasting Research Institute (FRI)
What the listing never says
- 19 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity
- No section describes what the person would actually do. Scope clarity
- No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency
- No location or timezone policy stated, so a candidate cannot tell where they may work from. Scope clarity
The listing, marked up
Job Description
• Architect, administer, and optimize Databricks workspaces, clusters, jobs, and workflows.
• Design and develop scalable Azure-based data engineering solutions, including Data Lake, Data Factory, Synapse, Key Vault, etc.
• Build high-performance data pipelines using Python and PySpark.
• Work with Hadoop-based platforms and distributed storage systems.
• Implement and manage real-time data streaming solutions and ingestion pipelines.
• Ensure proper configuration, monitoring, governance, and performance tuning of data platforms.
• Integrate CI/CD pipelines on Azure and work with containerization technologies (Docker/K8s).
• Develop and optimize SQL queries, stored procedures, and data models using MS SQL.
• Collaborate with architects, data scientists, and engineering teams to deliver enterprise-grade solutions.
• Define best practices, coding standards, and data engineering frameworks across the organization.
• Troubleshoot complex data issues and ensure high availability of mission-critical data systems.
Requirements
• Minimum 10+ years1 of total IT experience.
• Minimum 6+ years strong experience in the following areas: Databricks Administration , Azure Cloud , Python & PySpark , Hadoop ecosystem , Data streaming & data pipeline engineering , MS SQL
• Proven experience architecting large-scale distributed data systems.
• Strong understanding of data lakes, data warehousing, and big data architecture patterns.
Nice-to-Have Skills
• Experience with Apache Spark optimization and tuning.
• Hands-on experience with CI/CD pipelines on Azure (GitHub Actions, Azure DevOps, etc.).
• Knowledge of containers (Docker, Kubernetes).
• Understanding of MLOps or advanced analytics integrations.
Originally posted on Himalayas
• Architect, administer, and optimize Databricks workspaces, clusters, jobs, and workflows.
• Design and develop scalable Azure-based data engineering solutions, including Data Lake, Data Factory, Synapse, Key Vault, etc.
• Build high-performance data pipelines using Python and PySpark.
• Work with Hadoop-based platforms and distributed storage systems.
• Implement and manage real-time data streaming solutions and ingestion pipelines.
• Ensure proper configuration, monitoring, governance, and performance tuning of data platforms.
• Integrate CI/CD pipelines on Azure and work with containerization technologies (Docker/K8s).
• Develop and optimize SQL queries, stored procedures, and data models using MS SQL.
• Collaborate with architects, data scientists, and engineering teams to deliver enterprise-grade solutions.
• Define best practices, coding standards, and data engineering frameworks across the organization.
• Troubleshoot complex data issues and ensure high availability of mission-critical data systems.
Requirements
• Minimum 10+ years1 of total IT experience.
• Minimum 6+ years strong experience in the following areas: Databricks Administration , Azure Cloud , Python & PySpark , Hadoop ecosystem , Data streaming & data pipeline engineering , MS SQL
• Proven experience architecting large-scale distributed data systems.
• Strong understanding of data lakes, data warehousing, and big data architecture patterns.
Nice-to-Have Skills
• Experience with Apache Spark optimization and tuning.
• Hands-on experience with CI/CD pipelines on Azure (GitHub Actions, Azure DevOps, etc.).
• Knowledge of containers (Docker, Kubernetes).
• Understanding of MLOps or advanced analytics integrations.
Originally posted on Himalayas