Files
WeClone-new/weclone
xming521 5dcaad0e10 perf(infer): clear cuda cache after inference
Adds a call to empty the CUDA cache immediately after the LLM object is deleted, ensuring prompt release of GPU memory. This helps prevent out-of-memory issues and improves resource utilization for subsequent operations.
2025-08-17 15:18:37 +08:00
..
2025-07-01 20:00:34 +08:00
2025-04-21 20:30:55 +08:00