Showing posts with label VS1053b. Show all posts
Showing posts with label VS1053b. Show all posts

Sunday, May 2, 2021

ESP32-Radio (Part 1.5)

Introduction

It's been quite a long time since my last post and that is because I felt unreasonably tired lately, which might have been related to my work. In any case I've had a lot of progress with this project even if I was stuck with very basic, vary fundamental and very annoying issues for some time. Most of these issues are related to the basic things discussed in part one, so I'm writing this part "one and a half" before part two. I've also learned quite a few things related to the FreeRTOS that is included in the ESP-IDF. So let's discuss these few little things.

Buffer overflow

As discussed in part 1, the radio server seems to send data at a much faster rate than it's supposed to right after the connection has been established. I suspect this is done so that the client can fill the audio buffer quickly and start playing the audio. Without any check the audio buffer in ESP32 would fill up quickly and overflow while also making annoying noises. I fixed this issue by making the code loop until there is some free space in the buffer. This works "nicely" since the ESP32 has two cores and the other core can run the task that is responsible for sending the audio data to the decoder chip (among others).

At this point I think that is the worst way to solve this issue. The problem is that the code is simply stuck without the possibility to run any other code on that core. The correct way would be to use FreeRTOS Stream Buffers which would allow task switching in case the buffer is full. Using this would however require quite a lot of rewriting, which I didn't (yet) want to do. Eventually I realized that simply using vTaskDelay (in the loop) with a minimum time, would basically do the same thing. The code would check whether the buffer is full and issue this wait statement if it is. This in turn would allow other tasks to run on this same core instead of just running NOPs. Before this change the device definitely had some hiccups from time to time, but this change significantly improved performance and the issues are now gone.

VS1053b MIDI

As discussed in the first part, the VS1053b board from eBay has a design flaw, where the IOs are not pulled down as they are supposed to (as shown in the datasheet). This is related to the fact that this board is possibly the shittiest design that I have ever acquired from eBay. As explained here, "most modules will start in MIDI mode". There is also a link that points here. I originally failed to understand the fact that after making these "few register writes", it is necessary to make a soft reset so that the chip can read the state of the pins at boot.

Interestingly the device worked quite fine with my faulty code for a long while but then it simply stopped working. The funny part here is that when the device boots into MIDI mode, the AUDATA is configured to 44101 (which is okay for most servers and means 44100 Hz sampling rate and stereo mode) and the DREQ pin is high, which means that the device is constantly draining data and not playing anything.. However the latter link has a picture of pins that have to be physically connected (circled in green in the picture below) for the device to boot properly without any software fixes. I should have used the physical fix right from the start, but instead I wasted a lot of time making workarounds to detect the "stuck device" and rebooting it. That was a huge waste of time, cannot recommend. However here is the link to the original explanation right from the manufacturers of the chip.

VS1053b Board Buzzing

As I've explained in the first part, the board that I have, was making a very annoying and rather loud buzzing sound. The sound was coming from the board itself and couldn't be heard in the audio output. I spent quite some time poking around the VS1053b board without any luck. I also tried measuring voltages on the board, again without any luck. I tried to live with but after some time I was too annoyed again. It almost gave me headaches since it was so loud. Then I remembered how these things should be debugged: simply by taking a cap and poking it in different places. After some poking the buzzing suddenly stopped. After two months of ~constant buzzing it simply stopped (I had it connected to a USB port and thus it was on all the time my PC was on, just to see that it doesn't crash or disconnect).

And here's the explanation. The main output cap (C1) of the 1V8 regulator (U3, the core voltage) was located quite far away from the regulator. The regulator in question is AMS1117 and according to some datasheet it requires an insane 22 microfarad output capacitor. This capacitor should obviously be as close to the regulator as possible. Instead it's a few centimetres away and the connecting tracks are routed god knows where (shown in red and purple in the picture below). There is however a 100 nanofarad capacitor (C18) right next to the regulator, which, I suspect, was simply faulty in some way. After I've desoldered both 100n capacitors (3V3 and 1V8) and replaced them with 4.7 microfarad capacitors (C18 and C17 in the picture), the board has not made any sound (except for the audio output of course).


VS1053b evaluation board from eBay

I tried to understand this later on with the help of an oscilloscope without any luck. My final conclusion is that the original capacitor C18 in the picture above was faulty, since the board made no audible sound even without anything soldered in its place.

I've also acquired a second board to have just in case, and that board was even worse. It had the VS1003 chip instead of VS1053 and had a 2.5 V regulator instead of 1.8 V one. This part needs clarification. For the VS1053b the voltage ranges are 2.5 .. 3.6 V for AVDD, 1.7 .. 1.85 V for CVDD and 1.8 .. 3.6 V for IOVDD. So two regulators (3.3 V and 1.8 V) are enough for the VS1053b. For VS1003 however, the respective voltages are 2.6 .. 2.85 V for AVDD, 2.4 .. 2.85 for CVDD and CVDD-0.6 .. 3.6 V for IOVDD. Since the board looks the same as in the picture above, I assume that on this board AVDD is connected to 3.3 V. And the question is why would you do that? Absolute maximum ratings exist for a reason. Picture of the board is shown below.


VS1003 evaluation board from eBay

Perhaps less surprisingly this board didn't work correctly. It usually started okay and then the audio just faded out after some time of playing. I tried to debug it for a while and then simply gave up because the board is clearly not worth it.

Header issues

As I've explained in the first part, the server will send an HTTP header after the connection has been made. The new thing is that this header might vary a bit between different servers. This might be an issue, because I've only tried this with exactly two different servers and found three different issues. Here are the issues that I've encountered until now.

One server doesn't seem to send the sample rate of the audio at all. This is not an issue, since the sample rate can be read from the decoder chip and more specifically the AUDATA register. The value should be adjusted simply by (AUDATA >> 1) * 2 to get the actual value (to lose the stereo bit).

One server sends some garbage instead of the stream name. This is very annoying. However it can be easily solved by not using this value. The name of the radio station can simply be hard-coded with the URL address of the server in the device. I had an idea of fetching some channel list automagically for example from SHOUTcast API, but that's for another time.

One of the servers seems to use "content-type" as the tag while another uses "Content-Type". Solution for this is simply to use tolower function before trying to detect the tag.

Prototyping issues

This prototyping phase has lasted quite long for this project and the reason is that I want to try as much of the final implementation with it as possible. By as much as possible I mean the display and the buttons & rotary encoder. Since the pins of the ESP32 can be assigned quite freely, I can change the schematic based on the board design, to make the routing as simple as possible. The prototype can then be easily reconfigured to test the new connections before ordering the PCBs.

However there are a few issues related to the prototype. The worst issue is that because of the long wires, the device will easily pick up noise. I suppose it's not very dangerous, but suddenly hearing the audio at the maximum volume is certainly not fun. I've noticed this already at my desk and then I definitely noticed this when testing at the location where the end device is supposed to be used. Basically switching a noisy lamp or a power supply on or off will make the device output audio at full volume and also somehow garbled. Luckily the volume goes to the exact maximum, so it's very easy to detect as an error case.

I've concluded that this is not very good even for a prototype. So with my newly knowledge of FreeRTOS, I quickly made task that regularly checks whether the volume is at maximum and lowers it to some sensible level if it is. In retrospect, this fix had one of the best effect / effort ratios of all the things that I've made for this project.. Unfortunately the audio is still garbled so this fix doesn't solve that. I suppose it would require restarting the audio decoder. That would then require logic for finding the beginning of the next audio packet in the stream, or it would simply glitch for a while until the decoder gets the next header and recovers. In any case this should not be an issue with the final design and until then at least I don't need to hurt my ears.

Another issue related to the prototype is that some wires usually end up being near/above the WiFi antenna of the ESP32. This seems to be an issue because it introduces some noticeable noise in the audio output of the device. This is rather annoying because I can clearly hear it whenever there is a silent moment in the music. Interestingly there are no audio cables near the WiFi antenna, which means that the voltage regulation of the audio board is not very good. However it's enough to just insert some item between the antenna and the wires to get rid of this problem. For example a USB memory stick is very much enough. The signal strength will obviously be worse after this, but at least there is no noise in the audio whatsoever. This "fix" is okay for the prototype and should not be an issue with the final design. This is one good reason to follow the esp32 hardware design guidelines. :)

Final Words

I can only say that a lot of lessons were learned here. A simple RTFM can be applied here of course but it seems that FreeRTOS contains quite a few new things for me to learn. Also the issue with MIDI was just a very annoying mistake. The rest are simply adopting to the environment.

Sunday, January 17, 2021

ESP32-Radio (Part 1)

Introduction

Aim of this project is to make a physical, WiFi connected, Internet radio client (regardless of the slightly confusing name). Unlike many other projects, this is not my lifelong dream or anything, it's just that I'm very intrigued by connecting ESP32 to an audio decoder chip (VS1053b in this case) and by various possibilities these bring to the mix. I thought of this originally already when I was using an ESP8266 before ESP32 was released but I did not have enough knowledge back then and failed to google such project, so cannot really say if someone has already done such a thing then. Later on I had this idea again already with an ESP32 and after a quick search I found this mess. After skipping through the video without having the patience to watch it completely, I just decided to make one, although without the wire spaghetti and the Arduino crap.

Unlike other projects I've described here, I have a need to make this work as soon as possible. So this project could develop very fast (and have obvious issues) or it might never be finished. Who knows? However there's much to discover and learn here. I will start prototyping this project with eBay modules, so in the beginning it will be a mess of wires (like in the video) and will mostly be about the code, which I am making mostly from scratch, since I simply don't want to read 5k+ lines of code. There are also many new things here for me, both in software and in hardware, so this might get interesting.

I had some difficulties finding sources for this project (that don't involve Arduino) but here's one application note from microchip, here's one post in stackoverflow and one assignment(?). Also here is an Arduino library for VS1053b. Other sources are mentioned later on in this post. These were however a bit difficult to properly use in the text.

Background

Technically this project is very simple. An ESP32 is used to open a connection to a specific server which then sends audio data back to the ESP. The ESP should buffer this data to even out network latencies and stream it to the VS1053b decoder chip. The decoder chip then directly outputs analog signal which only needs to be filtered and possibly decoupled. The decoder chip is also able to output audio in a digital format, which would make it possible to connect it to a proper DAC if necessary.

SHOUTcast

The basic idea of SHOUTcast (or icecast) is basically to stream audio data. There's not much into it. First the server will transmit some kind of a header, which includes station name, genre, bit-rate and such. After that the server will start outputting audio data frame by frame. I'm not sure about the other formats, but at least MP3's contain a header in each frame, so the decoder chip can get all the needed information from the first few bytes of each frame. The server is also capable of sending metadata between chunks of audio data if the client requests so. The resulting data-stream would look as shown below assuming the metadata is requested. If the metadata is not requested, the stream only contains the header and continuous audio data.

[header] + [audio data] + [meta size] + [meta] + [audio data] + ...

1. The end of the header can be easily detected from the standard "\r\n\r\n" ending.

2. The size of the audio data chunk is specified in the header with the "meta-int" tag.

3. The "meta size" x 16 tells us how many chars there are in the metadata. Metadata is padded with zeros, if the length is not divisible by 16.

VS1053b

This is a very powerful chip that is able to decode MP3, AAC and OGG formats among others. The chip is also able to directly output analog signal as already mentioned and can drive headphones directly. The downside to this chip is its price, pitch of the pins and the amount of filter caps and routing it requires. However the good part is that it doesn't really need to be configured in any way. Basically it's enough to set the internal frequency, output volume and then just stream any supported audio data through the SPI bus. This makes the project very very simple. However there's much more to this chip, but that's for another time (but definitely for this same project).

Electronics

I will not provide any schematics here, since I simply improvised some connection between the ESP32 and the VS1053b. There is one SPI bus with two CS (Chip Select) pins, one for data and one for control. Additionally there is a reset signal and DREQ (Data REQuest). All this is exceptionally well documented in the datasheet. I used an ESP32-T development board, some VS1053b module from eBay (the oldest design I suppose, the one that doesn't look like an Arduino shield) and obviously DuPont cables.

A quick note here about the decoder module from eBay. It seems that the proper filter caps for the chip power are missing. Possibly for this reason the existing caps are making high pitched noise, which is quite loud and very very annoying. There is no noise in the audio output however. Additionally according to the datasheet the GPIOs of the chip should be connected to ground if not used, and they don't seem to be connected to ground in this board. This might be the reason that the chip does not automatically decode incoming MP3 audio and requires a few specific writes to do so.

I've also decided to build a prototype of the whole device using similar modules. However there is no such module for the display I want to use, so I had to design one myself. While at it, I also made a module that would fit a rotary encoder of my choosing and two buttons with pull-ups and filters. But that is for the next part.

Code

I've programmed this device using a lot of trial and error. The result can be split into five easy steps. This could be explained with many more steps, but I would like to strip down some things, like the ever changing API of ESP32. I only want to explain the general idea here.

Step 1, HTTP(S) GET

The first step after opening a connection to the server, is to make a HTTP GET request. It's a pretty standard request with the only exception of "Icy-MetaData:1" which tells the server to send the metadata amongst the audio data. The metadata is sent only if it has changed since the last time, so that it doesn't waste bandwidth. If the metadata has not changed since the last time, the server will still transmit the "meta size" byte, but it will be zero. Additionally there is an "Accept: */*" string which will tell the server that we accept any data. I suppose with this we could filter out formats that we don't support. Below is an example of the HTTP request, the WEB_PATH, WEB_SERVER and WEB_PORT have to be specified separately like in the http_request example. I have only tried the servers that use HTTP for now, but the same should work for ones that use HTTPS, just with different function calls.

static const char *REQUEST = "GET " WEB_PATH " HTTP/1.0\r\n"
    "Host: "WEB_SERVER":"WEB_PORT"\r\n"
    "User-Agent: esp-idf/1.0 esp32\r\n"
    "Accept: */*\r\n"
    "Icy-MetaData:1\r\n"
    "Connection: close\r\n"
    "\r\n";


Step 2, Parse header

The HTTP header can be used in this case to read the name of the stream (icy-name tag), genre (icy-genre tag), bit-rate (icy-br tag), audio format (Content-Type tag) and such. The bit-rate, sample-rate and audio format can be also read from the audio stream or the VS1053b chip after it starts decoding so it's not necessary to read these from the header. However it is necessary to read the icy-metaint value if the metadata has been requested. Other tags do not really affect functionality but are still good to fetch for example for displaying info on an LCD (like in the following parts of this project). Note: the Content-Type tag will say audio/mpeg for MP3 and audio/aacp for AAC.

Parsing the header is quite simple. The re-entrant strtok_r function can be used to first split the received data into lines and then another call can be used to split the tags and the payload. For testing these kind of things I like to use the C Playground. I just dump some example data from the device and try to parse it in the C Playground. It really speeds up the development process when I only need to test one piece of code and don't need to recompile the whole project and reprogram the ESP32. At least I'm not good enough to be able to write such code from the first try.

Speaking of testing, all the data can also be easily fetched on a PC using Unix tools. This could be done for example using wget as described here. I've personally also used curl. Both of these commands have to be terminated manually or they will keep loading data indefinitely. There is also a nice Unix tool called xxd that can be used to view hex data. Combined with less, xxd should be enough to inspect the stream data. Both of the commands also request the meta-data so that all that is written in this post can be easily verified.

wget --header="Icy-MetaData:1" -S -O reply.txt [url]
curl -H "Icy-MetaData:1" -i [url] --output reply.txt
xxd reply.txt | less


(Step 3, Parse metadata)

Technically audio data comes after the header, but in my opinion this step fits here better. Reading the metadata requires first waiting for "metaint" amount of bytes of audio data. After that, one byte should be read from the stream and multiplied by 16. The result is the amount of bytes that should be read as metadata (and not transferred to the decoder chip). After reading these bytes, the software should continue buffering audio data. An example of metadata is shown below.

StreamTitle='title of the song';

Parsing this syntax is quite annoying since the ' character can appear in the middle of the string. However I'm not sure if the ; character can appear in the payload. If it's not allowed, the whole string could be first split by the ; character and then = character, after which the first and the last character could be discarded.

Another issue here, at least at this level of programming, is UTF-8. Since we are making a simple embedded system, we do not want to support 1 112 064 different characters that the strings may contain. For this reason there should be at least some check that would simply replace the characters that our device cannot display with an underscore or any other character. I suppose for European radio stations it's enough to support Latin characters and the same but with all kind of hats (like åäöáàâ etc).

Step 4, Buffer audio data

Since there might be rather long latencies in the WiFi/Internet connection, some audio data should be buffered before forwarding it to the decoder chip. A simple circular buffer should suffice in this case. This source states that 300ms buffer should be enough, however I made a much longer buffer since there is plenty of memory in the ESP32.

Making a circular buffer is easy, it's just a large enough array with a pointer that wraps around to the beginning of the array after reaching the end. There should actually be two pointers, one for writing and one for reading. Special care should be taken so that the pointers don't pass each other, which would cause an ugly glitch in the resulting audio. It took me some time to implement this because of the way these two pointers are incremented.

The write pointer is incremented with whatever amount of bytes the receive function manages to receive from the Internet. That cannot really be specified, although I could make some function to write only until 32 byte boundary in the buffer. The read pointer is however always incremented by 32, because the datasheet specifies that at least 32 bytes can be written when DREQ goes high. For this reason I've implemented it in such a way that it does transmit exactly 32 bytes. Now this poses an issue, because we cannot check whether the read pointer is equal to write pointer, because there's 31/32 chance that the read pointer will skip write pointer and we will have a glitch. For this reason I've decided to make a check like so: read_ptr / 32 == write_ptr / 32. If the previous statement is true, the buffer is "empty" and that is technically an error (since the internal buffer of the decoder chip is quite small) and no audio data should be sent to the decoder (since there is no new data).

Originally I've expected the buffer never to overflow, since technically the decoder should consume the data at the same rate that the server sends it. More over I've assumed that the client (my device) should buffer some data first before sending it forward to the decoder. However that is apparently not quite so. It seems that the server sends data much faster right after opening the connection. I suppose this is so that the client can both fill the buffer and start playing immediately. This caused buffer overflows in my device, so I had to throttle the reception, because there's simply not enough memory for the amount of data some servers send. Making this functionality was easy, since the incoming data is processed byte by byte. The code was reduced to a simple "while((read_ptr + 1) % AUDIO_BUFFER_SIZE == write_ptr);" since reading is implemented in an interrupt and will not be prevented by this loop. This check should however be done before incrementing the pointer so that the device will not think that the buffer is empty.

Step 5, Transfer audio data to the decoder chip

As already mentioned, the easiest implementation is to transmit 32 bytes whenever the DREQ signal of the VS1053b goes high. However there might be large delays when receiving data from the Internet and the internal buffer of the decoder chip might not be enough for such a time. For this reason I've used a timer that calls an interrupt handler. This interrupt handler will then transmit audio data to the chip whenever the DREQ is high. This will interrupt any other ongoing tasks like the data processing, however it will not interrupt anything critical since those tasks should run on the other core and cannot be interrupted by the user program (AFAIK).

Making the transfer itself is quite simple. First the code should check whether DREQ is high. Then it should check whether there is at least 32 bytes of data in the buffer. If both checks pass, 32 bytes should be transmitted, the read pointer should be incremented by 32 and wrapped around if necessary.

Result

The result of this post is a completely hard-coded ESP32 Web Radio player that works mostly. It successfully connects to unsecured HTTP servers and streams data. It can also successfully strip and decode the header and metadata and print the results via serial line. The next step is to add a display, some inputs and make it configurable.

It feels amazing that less than 500 lines of code (excluding the WiFi and Internet connection code) is enough to make a hard coded Internet radio player. And most of this code is mostly written from scratch by me. However the length of the code might grow drastically as soon as the display is added and "hardcodeness" is removed. :)

Final words

This project is progressing quite rapidly. However I will need a breakout board to be able to prototype with the display that I want to use. And speaking of which, I only have one such display and more will be in stock in ... April apparently. I am not joking. This project might thus take quite a long time to be fully functional. I just hope I can have the first prototype PCB before my next vacation. :)

The most unfortunate thing about this project is that I've started it using the oldest ESP32 module that I had and there has been at least two new chip revisions after that. I don't really know if this affects anything in a meaningful way though. Additionally I also hope that this device will work in the location that it is designed for since I cannot often test it there. It might require an external WiFi antenna, but fortunately ESP32 modules with an external antenna are available.

Spoiler: I've been researching the VS1053b chip and there seems to be all kind of features including bass/treble control and plugin support that allows both an equalizer and a spectrum analyzer. With this it would be easy to make some kind of audio visualisation on a display or using external LEDs. In short, this decoder chip is very capable and I'm very interested in researching some of the available features.

Internet of Crosstrainers, part 2

Introduction As mentioned in the original post here , there were some issues in the described implementation. I had some difficulties to fin...