Machine Learning Engineer, Graph Deep Learning - Credit
Summary
Machine learning engineer building graph deep learning models (GAT/RGAT) over large-scale user relationship and financial behavior data to improve credit risk scoring. Day-to-day work covers training sample design, graph construction, model training and ablations, evaluation, and production deployment using Python, SQL, and PyTorch-based graph frameworks.
About the team:
Job Description:
Requirements:
We are looking for a machine learning engineer who is passionate about graph learning and solving real business problems. You will develop graph deep learning models using large-scale user relationships and financial behavior data to improve credit risk assessment.
You will contribute across the full development cycle, from training sample design and graph construction to model training, downstream evaluation, and production use. Our immediate focus is systematic experimentation with GAT and RGAT to deliver measurable incremental value to application scoring models, behavioral scoring models for existing customers, and other credit risk models. Over time, we will expand into heterogeneous graphs with multiple node types, temporal graphs, and graph self-supervised learning.
We welcome candidates from credit risk, e-commerce, recommendation systems, content integrity, social networks, and knowledge graphs. We value practical graph learning experience, rigorous experimentation, and the ability to translate research into business impact.
- Investigate training populations, historical windows, monthly snapshots, sampling and balancing strategies, repeated-user handling, and labels with different outcome horizons for A-card, B-card, and joint modeling. Assess their impact on performance and generalization.
- Design nodes, edges, and features from contact networks, transactions, shared devices, and other entity relationships. Experiment with individual and combined relation types, time windows, edge filtering, directionality, edge attributes, and high-degree node handling, balancing coverage, signal quality, and computational cost.
- Implement and improve graph attention models, including network depth, neighbor sampling, attention mechanisms, aggregation across relation types, and node feature selection. Use ablation studies to isolate contributions from a user's own features, neighborhood information, and graph structure.
- Compare pooled multi-snapshot training, continual training, and rolling retraining. Explore multi-task learning, regularization, and graph self-supervised learning to improve training stability, sample efficiency, and generalization over time.
- Incorporate graph scores or embeddings into existing risk models and measure incremental value over established features and baselines. Use AUC, KS, out-of-time (OOT) validation, segment analysis, and repeated experiments to assess performance, stability, and applicability across products and markets. Enforce point-in-time correctness for features, relationships, and labels to prevent leakage.
- Collaborate with data engineering, ML engineering, and risk modeling teams to improve data processing, graph construction, training, batch inference, and evaluation pipelines. Optimize runtime and resource usage, and support model versioning, deployment, and performance monitoring.
- Build on validated business use cases to explore heterogeneous graphs with user, device, merchant, and other node types; temporal graphs; and graph structure representations and link prediction for credit risk modeling.
- Bachelor's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, Data Science, or a related field.
- Strong foundations in machine learning and deep learning, including GNN message passing, neighborhood aggregation, and attention. Familiarity with models such as GAT, GraphSAGE, and GCN, with the ability to understand and implement approaches for graphs with multiple relation types.
- Proficiency in Python, SQL, and PyTorch, with practical experience in DGL, PyTorch Geometric, or a comparable graph learning framework. Ability to process data, implement models, debug training, and analyze results.
- At least one graph learning research, internship, or industry project, such as node classification, graph representation learning, heterogeneous graph modeling, knowledge graph completion, or link prediction. Ability to explain graph construction, modeling choices, experimental design, and results clearly.
- A rigorous approach to experimentation and validation, including appropriate baselines and ablations, with attention to data quality, temporal splits, information leakage, and model stability.
- Ability to read technical papers and communicate technical ideas effectively, together with strong learning, problem-solving, and collaboration skills.
Good to have:
- Experience in credit risk, fraud detection, or fintech modeling, including A-card / B-card models, observation and performance windows, label maturity, and risk model validation.
- Experience in heterogeneous graph node classification, multi-relational modeling, knowledge graph completion, inductive link prediction, or relational reasoning, with the ability to adapt these methods to new business settings.
- Experience in temporal graphs, graph self-supervised learning, graph pretraining, or multi-task learning.
- Experience processing large-scale graph data, using Spark / Hive, optimizing neighbor sampling or GPU training, or working with distributed training.
- High-quality publications, open-source contributions, or demonstrated business impact in graph learning, representation learning, or related areas.