Train: Optimised single_pass — charbonnier ramp + sky0.7 area-pooled + 50% full-noise (C2) 9ed518c3
Supersedes config 93 (“L1@10 + sky0.7 + 50% full-noise (C)”, run 2026-07-28_10-02-59_4779) after a conceptual review of commit 3cbdf26. Keeps the intent of (C) — supervised sky at 0.7 and 50% of examples at the zero-terminal-SNR endpoint — and fixes four issues that made it uninterpretable or actively harmful. 1. Prediction loss: the hard L2→L1 step at epoch 10 is replaced by an L2→Charbonnier ramp over epochs 10–14. A hard swap roughly doubles both primary losses in one epoch (E|d| = sqrt(2/pi)*sqrt(E[d^2])), shifting the primary/auxiliary balance and gradient clipping at the same instant; pure L1 also has non-vanishing gradient at the optimum, which leaves a dither floor — the opposite of the local-structure goal. Charbonnier is quadratic near zero and linear on large residuals. 2. LR schedule: constant_with_warmup → cosine_with_warmup. A constant LR is what makes the L1-family gradient floor bite; the template's own comment already recommends cosine here. 3. Latent weight pooling: new loss.depth_weight_pooling=area. With min pooling a latent cell mixing valid depth and sky collapses to the sky weight, so 0.7 attenuated exactly the thin structures against sky that this experiment targets. Area pooling keeps the all-or-nothing gate on genuinely missing pixels. 4. VKITTI2 valid_max_depth 300 → 650. VKITTI2 stores 16-bit centimetre depth, so its sky sentinel is 655.35 m; at 300 m the above_max fallback relabelled all real geometry between 300 m and the sentinel as sky and, at sky weight 0.7, actively taught it the 800 m override. Also enables depth.neutralize_invalid_depth on both training sources: under log_depth, clamping mapped 0/negative/-inf pixels onto the NEAR plane (1 m), and the VAE encoded that into neighbouring supervised latents despite zero loss weight. RDS retains valid_max_depth == sky_depth_override == 800, so its above_max sky mask may be empty. This is left unchanged deliberately because the raw RDS sky sentinel was not verified; training now warns at startup and logs realized per-source mask coverage each epoch. Check the rds `sky` fraction on epoch 0 before comparing sources. Requires the reviewed fixes on the training branch (depth_weight_pooling, charbonnier registry entry, neutralize_invalid_depth, all-breakpoint schedule validation, sky-inclusive smoothness region, and the depth_sky/nonsky prediction-loss split).
Output
Failure Metadata
failedContext JSON
{
"reason": null,
"exitCode": 1,
"remotePid": 1192456,
"computePid": null,
"allocationJobId": null
}Datasets
14clear_root_rds/cluster/work/igp_psr/drothenpiele/data/rds/rgb.zip/depth_root_rds/cluster/work/igp_psr/drothenpiele/data/rds/depth.zip/intrinsics_root_rds/cluster/work/igp_psr/drothenpiele/data/rds/calibration.zip/train_hazy_root_rds/cluster/work/igp_psr/drothenpiele/data/rds/gloomy_fixed/foggy_rgb.zip/val_hazy_root_rds/cluster/work/igp_psr/drothenpiele/data/rds/gloomy_fixed/foggy_rgb.zip/clear_root_vkitti2/cluster/work/igp_psr/drothenpiele/data/vkitti_2.0.3_rgb.zipdepth_root_vkitti2/cluster/work/igp_psr/drothenpiele/data/vkitti_2.0.3_depth.zipintrinsics_root_vkitti2/cluster/work/igp_psr/drothenpiele/data/vkitti_2.0.3_textgt.zip/train_hazy_vkitti/cluster/work/igp_psr/drothenpiele/data/vkitti_next/gpu_2/foggy_rgb.zip/val_hazy_vkitti/cluster/work/igp_psr/drothenpiele/data/vkitti_next/gpu_2/foggy_rgb.zip/muses_rgb/cluster/work/igp_psr/drothenpiele/data/muses/frame_camera_trainvaltest.zip/muses_sparse_depth/cluster/work/igp_psr/drothenpiele/data/muses/lidar_trainvaltest.zip/muses_intrinsics/cluster/work/igp_psr/drothenpiele/data/muses/frame_camera_trainvaltest.zip/muses_extrinsics/cluster/work/igp_psr/drothenpiele/data/muses/frame_camera_trainvaltest.zip/Execution Artifacts
2Published Exports
Launch-Owned Exports
These semantic handles are what downstream pipelines should resolve against. Auto-published exports come from config-template metadata captured on this launch.
This launch does not publish any exports yet.
Parameters
278Typed Parameters
clear_root_rds(4)depth_root_rds(5)intrinsics_root_rds(6)train_hazy_root_rds(195)val_hazy_root_rds(196)clear_root_vkitti2(1)depth_root_vkitti2(2)intrinsics_root_vkitti2(12)train_hazy_vkitti(216)val_hazy_vkitti(238)muses_rgb(77)muses_sparse_depth(76)muses_intrinsics(82)muses_extrinsics(81)Simple Parameters
cliptruecropcenterseed42trackerwandbuse_emafalsegpu_typertx_4090:1job_namesd_train_concatnorm_max1norm_min-1run_namejoint-opt-charbonnier-sky-area-noisetmp_size10Gadam_8bittruedata_kindreal_drive_simema_decay0.9999ema_dtypefloat32log_every10max_depth800min_depth1adam_beta10.9adam_beta20.999batch_size16clear_root4depth_modemetric_logdepth_root5image_size[384, 768]joint_modesingle_passlog_imagestruenum_epochs200output_dir/cluster/scratch/drothenpiele/SD21/exp_1time_limit2-00:00:00mem_per_cpu8Gnum_workers4adam_epsilon1e-8adam_foreachfalseaspect_ratio[1, 2]camera_modelpinholeconditioningconcatlambda_depth1lambda_point0.0min_lr_ratio0.01project_namejoint-dehazinguse_geometryfalsewarmup_ratio0.05weight_decay0cpus_per_task8lambda_camera0.0lambda_dehaze1lambda_eg_ssi0.0learning_rate0.00004max_depth_rds800max_grad_norm1min_depth_rds0.00001val_hazy_root104camera_enabledtruefreeze_encoderfalselog_visibilitytruenum_log_images4use_depth_headtrueval_batch_size10camera_fallbackinvaliddepth_head_typebase_residualenable_xformerstrueeuler_train_dir/cluster/work/igp_psr/drothenpiele/data/out/train/sd21-jointlambda_depth_tv0.0002mixed_precisionfp16prediction_typev_predictionsource_kind_rdsreal_drive_simtrain_hazy_root103visibility_cmapviridiscamera_embeddingsinecamera_emit_raysfalsecamera_mlp_ratio4camera_num_heads4camera_token_dim256depth_noise_typeannealed_multireslambda_log_depth0.0lr_schedule_typecosine_with_warmupmin_max_quantile0.02no_decay_enabledfalsepretrained_modelsd2-community/stable-diffusion-2-1sky_depth_meters800camera_num_layers2camera_num_tokens4conv_in_init_modejoint_marigoldgeometry_lambda_z0.15geometry_timestep0joint_weight_psnr1joint_weight_ssim10lambda_depth_head0.2lambda_lambda_mse0.0lambda_visibility0.05max_depth_vkitti2650min_depth_vkitti20.00001sampling_strategysource_balancedsource_weight_rds1depth_ensemble_tol0.001lambda_uncertainty0.0model_depth_outputvae_decodeval_every_n_epochs1camera_feature_poolavgdepth_ensemble_size1geometry_lambda_phi1num_inference_steps25rds_val_max_samples128sampling_epoch_sizeallsave_every_n_epochs2source_kind_vkitti2vkitti2use_visibility_headtruewarmup_start_factor0.001camera_metadata_rootnullconv_in_target_scale0.7071external_camera_modenonegeometry_loss_robustcharbonnierlambda_vae_depth_rec0.0lambda_visibility_tv0.005sky_depth_meters_rds800camera_branch_enabledtruecamera_embedding_sizelatentdepth_gradient_robustcharbonnierdepth_multires_levels4encoder_learning_rate0.000007geometry_head_enabledfalsegeometry_lambda_theta1lambda_depth_gradient0.05lambda_geometry_total0.0source_weight_vkitti21use_cross_task_fusiontruevae_decoder_trainablefalsevisibility_rank_pairs128weight_decay_backbone0weight_decay_no_decay0zero_grad_set_to_nonetruecamera_feature_sourcesautocamera_metadata_formateuler_loadingdepth_ensemble_max_res1024depth_smoothness_spacelatent_x0freeze_cross_attentiontruegradient_checkpointingtrueinference_depth_outputvae_decodelambda_geo_consistency0.0lambda_visibility_rank0use_task_skip_adaptersfalsevisibility_rank_margin0.05camera_embedding_stride8camera_prediction_spaceresidual_intrinsicsdepth_ensemble_max_iter2depth_multires_strength0.9depth_smoothness_scales4geometric_pairs_enabledfalsejoint_noise_correlation0joint_weight_delta1_pct0.5keep_last_n_checkpoints5validation_depth_outputvae_decodevisibility_target_gamma4visibility_warmup_steps1000vkitti2_val_max_samples0cross_task_fusion_blocks[0,1,2,3]cross_task_fusion_detachfalsecross_task_fusion_kerneladaptivedepth_ensemble_reductionmediandepth_normalization_typelog_depthgeometry_camera_encodingsinegeometry_head_num_blocks2joint_weight_abs_rel_pct0.5sky_depth_meters_vkitti2800camera_intrinsics_loadinghierarchicaldepth_diagnostics_enabledtruedepth_head_residual_scale0.05geometry_depth_conventionz_depthgeometry_metric_depth_max800geometry_metric_depth_min1num_inference_steps_depth20vae_decoder_learning_rate0.000001visibility_apply_to_depthfalsecamera_residual_logit_clip4depth_head_hidden_channels64depth_resize_interpolationnearestdepth_smoothness_normalizetruedetach_rays_for_point_losstrueenable_efficient_attentiontruelambda_visibility_preserve0.05visibility_apply_to_dehazetruevisibility_apply_to_fusionfalsevisibility_hidden_channels32visibility_target_quantile0.9weight_decay_depth_decoder0weight_decay_geometry_head0camera_require_for_trainingtruecamera_sine_num_frequenciesnullcheckpoint_selection_metricvalset/muses/maecross_task_fusion_dilations[1,2,4]depth_decoder_learning_rate0.00002emit_camera_full_resolutiontruegeometry_eg_ssi_edge_sourcehazygeometry_encoder_input_modehazy_hazygeometry_head_learning_rate0.00004geometry_output_uncertaintytruegeometry_ray_representationunitgradient_accumulation_steps1validation_geometry_enabledfalseweight_decay_dehaze_decoder0camera_output_representationanglescamera_valid_for_approximatefalsecross_task_fusion_directions["rgb_to_depth","depth_to_rgb"]dehaze_decoder_learning_rate0.00002depth_head_context_dilations[1, 2, 4]depth_smoothness_edge_sourcecleardepth_smoothness_edge_weight10geometry_head_residual_scale0.05lambda_depth_edge_smoothness0.002lambda_rgb_depth_consistency0.1weight_decay_geometry_camera0depth_diagnostics_edge_sourcedehazeddepth_diagnostics_edge_weight10depth_head_base_upsample_modebilineargeometry_camera_learning_rate0.00004geometry_head_hidden_channels128geometry_sine_num_frequencies64visibility_decoder_grad_scale1.0camera_approximate_fov_degrees60checkpoint_selection_directionmincross_task_fusion_learned_gatetruegeometry_use_deterministic_vaetruenative_crop_lower_region_ratio0.05depth_multires_downscale_factor2lambda_depth_gt_edge_smoothness0validation_geometry_camera_modepredictedvalidation_geometry_output_modegeometry_headcamera_include_input_in_encodingnullcamera_intrinsics_metadata_scopeintrinsicscross_task_fusion_gate_init_bias0cross_task_fusion_sender_lowpassfalsedepth_head_use_depthwise_contexttruegeometric_pairs_views_per_sample2geometry_condition_on_visibilityfalsegeometry_detach_encoder_featurestruecross_task_fusion_hidden_channels64depth_diagnostics_normalize_depthtruenative_crop_centered_region_ratio0.0native_crop_lower_region_positioncenter_bottomrgb_depth_consistency_edge_weight4rgb_depth_consistency_weight_modehallucinatedgeometric_pairs_emit_pair_metadatatruegeometry_include_input_in_encodingfalselambda_depth_multiscale_smoothness0.001depth_ensemble_regularizer_strength0.02native_crop_centered_region_positioncentercross_task_fusion_attention_kv_tokens1024cross_task_fusion_attention_max_block-1depth_diagnostics_road_edge_threshold0.08geometry_head_detach_camera_embeddingtruedepth_diagnostics_road_bottom_fraction0.4geometry_encoder_no_grad_when_detachedtruecross_task_fusion_sender_lowpass_kernel3depth_diagnostics_high_frequency_kernel5validation_geometry_compute_edge_metricstruevalidation_geometry_compute_point_metricsfalsevalidation_geometry_max_points_per_sample20000validation_geometry_save_geometry_samplestruevalidation_geometry_compute_camera_metricstruevalidation_geometry_compute_resolution_sweepfalsevalidation_geometry_max_point_metric_samples8validation_geometry_save_point_cloud_samplesfalsevalidation_geometry_compute_uncertainty_metricstrueRaw JSON
{
"clip": "true",
"crop": "center",
"seed": 42,
"tracker": "wandb",
"use_ema": "false",
"gpu_type": "rtx_4090:1",
"job_name": "sd_train_concat",
"norm_max": 1,
"norm_min": -1,
"run_name": "joint-opt-charbonnier-sky-area-...