DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
New Features
- Added Scala Inference APIs (#9678). See: MXNet Scala Inference API
- Added module to import ONNX models into MXNet (#9963). See: Proposal: ImportExport module
- Added support for Model Quantization with Calibration (#9552).
- Added Exception Handling support for operators and iterators (#9681). See: Improved Exception Handling in MXNet
- Added MKLDNN support for MXNet (#9677). See: MKLDNN integration
- Added FP16 support for distributed training (#10183).
- Added Sparse support for Custom Operator (#10374).
- Add multi-proposal operator (CPU version) and fix the bug in multi-proposal operator (GPU version) (#9939).
- Profiling enhancements - VTune objects, individual operator profiling, C API profiling, Memory usage profiling profiling (#8972)
Bug-fixes
- Test fixes Fixed tests - Flakiness/Bugs - (#9598, #9951, #10259, #10197, #10136, #10422). Please see: https://github.com/apache/incubator-mxnet/projects/9
- Fixed crash when profiler not enabled (#10306).Fixed uncaught exception for bucketing module when symbol name not specified (#10094).
- Fixed regression output layers (#9848).
- Fixed crash with mx.nd.ones (#10014).
- Fixed sample_multinomial crash when get_prob=True (#10413).
- Fixed buggy type inference in correlation (#10135).
- Fixed race condition for CPUSharedStorageManager->Free and launched workers at iter init stage to avoid frequent relaunch (#10096).
- Fixed DLTensor Conversion for int64 (#10083).
- Fixes for profiler (#9932, #10306)
- Fixed ndarray assignment issues (#10022, #9981).
- Fixed incorrect indices generated by device row sparse pull (#9887).
- Fixed print_summary bug in visualization module (#9492).
- Fixed cast storage support for same stypes (#10400).
...
Performance Improvements
- Improve sparse.adam_update (Improved sparse SGD, sparse AdaGrad and sparse Adam optimizer speed on GPU by 30x (#9561, #10312, #10293, #10062).
- Improve sparse sgd on GPU (#10293).
- Improved 'sparse.retain' performance on CPU by 2.5x (#9722)
- Replaced Replace std::swap_ranges with memcpy (#10351)
- Implement Implemented DepthwiseConv2dBackwardFilterKernel from tensorflow codebase, which is over 5x faster (#10098)
- Implemented CPU LSTM Inference (#9977)
- Added Layer Normalization in C++ (#10029)
- Optimized Performance optimized for rtc (#10018)
- Added Parallelization for ROIpooling OP (#9958)
- Accelerate Accelerated the calculation of F1 (#9833)
...