Startup Spotlight: Closing The Gap From Design To Code: Transforming Designs Into Code
Checking session availability…
Hang tight while we load the latest updates.
- Why we need to rethink the current front-end development workflow
- How recent breakthroughs in AI are going to transform the way we build software
- What technology is behind the pix2code prototype?
- How will UIzard allow front-end developers to focus on what matters?
Startup Spotlight: Closing The Gap From Design To Code: Transforming Designs Into Code
Tony Beltramelli at UXDX EMEA. Video: https://youtu.be/OpwOBVPAdlY
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Transforming your design into code
[00:00:00] I'm going to talk about transforming your design into code. What I mean by that is, right now, as maybe some of you may know, when you want to build user-facing applications you'll typically involve designers who will be in charge of creating all the screens, the pages, the layout, and then they will ship this to a front-end developer who will turn this into code, namely HTML and CSS for web apps, XML for Android and so on. And then front-end developers will really shine when they implement all the functionality, the features and the logic of the application.
[00:00:30] But as some of you may know, if you've been in this spot, building interfaces as a developer is quite a frustrating experience. It just kills the iteration cycle with the designers, it is expensive, and if you ask any front-end developer to describe the experience of having to write HTML and CSS they will typically describe something along these lines.
[00:00:52] And to give you a sense of scale, every single day about twelve thousand websites are being uploaded online, and about a thousand apps on the different marketplaces, which means a lot of frustration is going on as we speak.
[00:01:12] So the idea is, to solve this problem, well, we can just build a code generator. But the problem is, for which graphics software should you build the code generator? Some designers are using Photoshop, some are using Sketch, and there will probably be some new hot software in two months from now that we don't expect. So how can you build a code generator to match those different tools?
[00:01:38] And the other problem is, if you are about to generate code, which platforms should you support, which library should you include? I like tabs, my colleagues prefer spaces. I want my code to look good for me. How do we do that? How do we fit every single developer's needs? A solution, as I said, should work with any graphical editor and it should produce code which is customized for everyone.
[00:02:07] So the idea is: can we use AI, which is kind of a buzzword this year, to automate part of this process? Can we somehow put an intelligent algorithm in between that would generate the code and allow the developer to focus on what matters, which is implementing the logic, implementing the features? This is basically what we did.
Putting the jargon aside: why machine learning
[00:02:34] Just to put some terms and jargon away: AI is just a research field, within which there is machine learning, which is another research field. But when everybody is talking about AI right now they talk about deep learning. Deep learning is really the rock star of AI right now, and 99% of all the different research papers are using deep learning. So when I talk about AI, I am talking about deep learning.
[00:02:57] So why do we need machine learning? Couldn't we just generate code the old good way, using software engineering and an engineering pipeline? Well, the thing is, without AI, or using traditional AI as I put it in quotes, you will have to pay experts to manually input rules and knowledge into the algorithm.
[00:03:20] So let's say you want to build an algorithm to detect cat pictures. Then you will hire a bunch of cat experts who will write down programming rules about what a cat is. Which has limitations, because if tomorrow you try to detect a cat which is pink, and none of your experts thought about encoding the concept of a pink cat, then it won't work.
[00:03:44] Using machine learning, you input a lot of training examples, and instead of providing the algorithm with rules and knowledge you will just tell the algorithm what the end result should look like. So if you want to detect cat pictures you'll just provide a lot of different cat pictures and tell the algorithm, this is a cat. And eventually, if there is a pink cat showing up, your algorithm should be able to understand the concept of a cat and understand cats.
[00:04:09] So using machine learning and having this very nice generalization feature, can we actually build an AI to generate code for user interfaces? Wouldn't it be cool if you could just upload a user interface image, which can be produced with any software, have some magic happening in the cloud, and download the code back?
[00:04:35] What I'm going to describe now is pretty much a prototype to do just this. I want to describe the way we did it in order to inspire you to maybe use machine learning in your own product, for your own problems, and also to let you realize that having a machine learning, AI-powered solution is actually not that hard. Which means we're going to use AI in our software engineering process in a very soon, very small time scale.
Borrowing from image captioning
[00:05:02] If we look at the problem a little bit more carefully, what we are trying to do is to automatically identify graphical user interface components, such as a button, and generate a code snippet that describes this user interface component. And you want to do this automatically, so the computer only sees pixel values from 0 to 255. Sounds like a pretty tough problem.
[00:05:25] But for every problem, where we really search for inspiration is in research, and what other people are doing. Deep learning has been extremely good recently at doing image captioning. Image captioning is the task of generating an English description given a photograph. This is basically what Facebook is using to describe pictures to blind people. So it works, it's not science fiction any more, this is like two-year-old research.
[00:05:53] And the assumption was: if you can use machine learning to generate an English description given a photograph, can you use the same approach to generate code given a graphical user interface mockup? In both cases you are trying to generate text from an image.
[00:06:11] And the nice thing about deep learning is that it's a statistical environment[?]. If you want to detect a dog or a cat or whatever, it doesn't matter which color the dog is, the algorithm should understand the concept of a dog. And the same with UI: a button can take different shapes, different colors and so on. And once again, for the concept of a dog it doesn't matter where the word dog is in a sentence, and the same applies to code: whatever the place in your code where the button snippet should be, it's a button regardless of the position.
[00:06:47] And if you cut the problem into some small pieces, you basically have a computer vision problem of understanding the scene, and deep learning gives us a really nice tool to do just that, which is the convolutional neural network. I'm not going too much into detail about what it is, but on a higher level a convolutional neural network allows you to extract features from images without having to tell the algorithm what to do. So if you want to build an image classifier that recognizes a hot dog, you will just feed it a bunch of hot dog pictures, and eventually the convnet will be able to understand what is a hot dog and what is not a hot dog.
[00:07:25] And the other problem, if you want to generate code, English or any sequence of data: to understand language you need to be able to have an algorithm that understands the order of words, the order of the tokens. And deep learning research gave us a really nice tool to do just this, which is recurrent networks. Once again, without going into too much detail, recurrent nets have this recursion which allows you to model sequence data. This is extremely useful if you want to predict time series data, as people are using it to break the trading market or the stock. But you can also consider text as a sequence of characters, and code as a sequence of tokens.
A domain-specific language in the middle
[00:08:06] So now that we are aware of those two different components, we can assemble them as Lego blocks and try to build a machine learning algorithm that generates code from user interface images. Now, what output should we have? We already know that we're going to use an image as input, but do you want to output HTML, do you want to output CSS, do you want XML?
[00:08:32] We basically solved this problem by introducing our own language. This very simple domain-specific language is designed to describe what a user interface is made of, regardless of you working with mobile interfaces, desktop interfaces or web interfaces. All of these different interfaces contain buttons, labels, sliders and so on. So we built this very simple programming language, and then we can train our AI to understand this language alone and compile from this language to HTML, to XML and so on. So we have control over the target.
[00:09:10] So this is the model. I'm not going too much into detail because time is running. And then when you generate the code you basically generate one token at a time. You provide the graphical user interface and you tell the algorithm, start generating tokens, and after some time it generates one token at a time and you get the full code generated for a given user interface.
[00:09:36] And the very nice thing about this very simple architecture is that you use an image as input, so you could imagine inputting graphical user interfaces produced with any software. If some of your designers want to use Microsoft Paint because it's pretty cool, then in theory you could use the same approach using Microsoft Paint user interfaces. And on the developer side, because we have this domain-specific language, we can compile it to anything. If tomorrow Google or Facebook or whoever releases a brand new front-end library, then we can support it. We just have to build a compiler and we can use exactly the same AI to generate the code. Pretty cool.
Does it work
[00:10:20] Does it work? Well, our proof of concept actually worked, to our greatest surprise. Here there is an example with a very simple iOS interface, and as you can see here on the left, this is the expected result, and on the right this is the AI-generated user interface. There are some mistakes but it looks pretty close.
[00:10:42] Here on an Android user interface you can see there are some mistakes: for whatever reason the AI couldn't generate some of the buttons. Just to show you where the pitfalls are at the moment. And here, interestingly, on the Bootstrap web interface, for whatever reason the AI could model the layout of the screen, but for whatever reason it didn't care about the color of the buttons.
[00:11:06] The stack is built on Python using the TensorFlow and Keras deep learning libraries, and everything is trained on GPU on AWS. For those of you who want to play with this proof of concept, the code is available on GitHub, datasets as well, and there is also a demo video for you guys to check out the current version of the algorithm.
Training data and the future workflow
[00:11:30] So, what now? We've heard about deep learning a lot during the last year. It's being applied to a lot of different industries, but one of the main problems in using deep learning for real-world problems is the access to training data. And the reason I'm extremely excited about applying deep learning to front-end development is that there is an amazing amount of data. If you look at the web and you build a web crawler, you could imagine collecting screenshots of websites with the HTML code, and you basically end up with an infinite amount of training data. So it's really just a matter of time before front-end developers only have to focus on the logic and the UI part is completely automated.
[00:12:12] The other thing is that you don't need to have manually labeled data. If you want to build a hot dog classifier AI, you will have to pay hundreds of hot dog experts to label the data as hot dog or not hot dog pictures. In our case, for any given user interface screenshot, the label is the code. So there is no problem generating training sets.
[00:12:39] So I wanted to think a little bit about what the future workflow should look like, because if you have an AI which can directly generate HTML and CSS and then let developers focus on the logic, the amount of iteration cycles you can have with the designers is just boosted. You can imagine having way better products because you have more iterations in the same amount of time.
[00:13:02] And if you think a little bit further: if you have an AI which can understand the concept of graphical user interfaces just with an image, you could basically let anyone use AI to build applications. Anyone is able to take a piece of paper, sketch a user interface, then you throw this to the AI and the AI can generate design files or code. So the idea is that we are really on a transition model right now where anyone could potentially become an application builder without any knowledge of code or design in the first place.
[00:13:37] So we really believe in a future where humans and robots collaborate to build better products, hopefully. And that's about it. I would be happy to answer your questions. Thank you.

