[Remote] Data Engineer (Python, PySpark, Databricks)

Note: The job is a remote job and is open to candidates in USA. GSquared Group is seeking a hands-on Data Engineer with strong Python and PySpark development experience in a Databricks environment. The role involves designing, developing, and maintaining scalable data pipelines while collaborating with various teams to deliver high-quality data solutions.


Responsibilities

  • Design, develop, and maintain scalable data pipelines using Python and PySpark
  • Build and optimize batch and streaming data pipelines in a Databricks environment
  • Develop reusable, production-ready code to ingest, transform, and load large-scale datasets
  • Implement data quality validation, error handling, monitoring, and alerting to ensure reliable data processing
  • Optimize Spark jobs for performance, scalability, and cost efficiency
  • Collaborate with data architects, analytics teams, and business stakeholders to deliver high-quality data solutions
  • Participate in code reviews and contribute to engineering best practices
  • Deploy and support data pipelines through CI/CD processes and automated release pipelines
  • Troubleshoot and resolve production issues while continuously improving pipeline reliability and performance
  • Contribute to the design and evolution of modern cloud-based data platforms

Skills

  • 5+ years of professional Data Engineering experience
  • Strong hands-on experience developing data pipelines using Python and PySpark
  • Experience working with Databricks
  • Experience building and supporting production-grade data pipelines processing large datasets
  • Strong understanding of distributed data processing concepts and Spark optimization techniques
  • Experience implementing data quality, validation, monitoring, and logging practices
  • Experience with orchestration tools such as Databricks Workflows, Apache Airflow, or similar
  • Experience working with Git and modern CI/CD deployment practices
  • Strong SQL skills and experience developing efficient data transformations
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Excellent troubleshooting, analytical, and problem-solving skills
  • Experience with Delta Lake and the Databricks Lakehouse Platform
  • Experience building metadata-driven or reusable pipeline frameworks
  • Knowledge of streaming technologies such as Apache Kafka or Spark Structured Streaming
  • Familiarity with modern data architecture patterns, including Medallion Architecture

Company Overview

  • GSquared Group provides information technology staffing and recruiting services. It was founded in 2010, and is headquartered in Alpharetta, Georgia, USA, with a workforce of 11-50 employees. Its website is https://gsquaredgroup.com/.

  • Share Share
    Apply Now →