No
Yes
View More
View Less
Working...
Close
OK
Cancel
Confirm
System Message
Delete
My Schedule
An unknown error has occurred and your request could not be completed. Please contact support.
Scheduled
Scheduled
Wait Listed
Personal Calendar
Speaking
Conference Event
Meeting
Interest
There aren't any available sessions at this time.
Conflict Found
This session is already scheduled at another time. Would you like to...
Loading...
Please enter a maximum of {0} characters.
{0} remaining of {1} character maximum.
Please enter a maximum of {0} words.
{0} remaining of {1} word maximum.
must be 50 characters or less.
must be 40 characters or less.
Session Summary
We were unable to load the map image.
This has not yet been assigned to a map.
Search Catalog
Reply
Replies ()
Search
New Post
Microblog
Microblog Thread
Post Reply
Post
Your session timed out.
This web page is not optimized for viewing on a mobile device.Visit this site in a desktop browser or download the mobile app to access the full set of features.
2018 GTC San Jose
Favorite
Remove from My Interests

S8576 - Achieving Human Parity in Conversational Speech Recognition Using CNTK and a GPU Farm

Session Speakers
Session Description

Microsoft's speech recognition research system has recently achieved a milestone by matching professional human transcribers in how accurately it transcribes natural conversations, as measured by government benchmark tasks. In this talk we will discuss the significance of the result, give a high-level overview of the deep learning and other machine learning techniques used, and detail the software techniques used. A key enabling factor was the use of CNTK, the Microsoft Cognitive Toolkit, which allowed us to train hundreds of acoustic models during development, using a farm of GPU servers and parallelized training. Model training was parallelized on GPU host machines, using 1-bit distributed stochastic gradient descent algorithm. LSTM acoustic and language model training takes advantage of CNTK's optimizations for recurrent models, such as operation fusion, dynamic unrolling, and automatic packing and padding of variable length sequences. We also give an overview of CNTK's functional API.


Additional Information
Speech and Language Processing
Higher Education / Research, Software
Intermediate technical
Talk
50 minutes
Session Schedule