INFERENCE ALLOCATION FLOW PULL LOAD API Gateway Registry Allocator GPU 1 GPU 2 GPU 3 GPU 4 TENSOR_PARALLEL_SIZE=4